Find published dossiers by topic, company, product or technology.

Showing 121–140 of 296 dossiers

Meta Muse turns a consumer AI assistant into a persistent agent — and its first Mac zero-day tests the containment model

Muse packages persistent autonomous execution, credentials, payments, app access and memory into a mainstream consumer product. A September macOS hotfix now provides an early real-world lesson: agent containment has to protect not only the cloud runtime but also the local control path into the agent.

Cloudflare Worker Previews gives every Git branch its own isolated runtime

Branch previews are common for frontend code, but Worker Previews extends the boundary to the runtime itself. Each branch can have independent bindings, state and logs, making parallel human and agent work safer while preserving a production-like execution path.

Claude Code Projects turns one engineering goal into parallel cloud-agent branches

The important change is not simply that Claude can run several agents. Projects now owns decomposition, shared context, branch isolation and progress coordination across full Claude Code sessions, while the trade-offs become usage burn, cloud-only execution and ordinary merge conflicts when parallel work overlaps.

GitHub rewrote Copilot’s 800,000-line agent runtime in Rust with agents doing most of the coding

The useful part is not the 800,000-line headline. GitHub has published unusually detailed receipts for a production-scale agent-assisted migration: roughly $120,000 of token spend, 14.5 weeks of incremental releases, dozens of regressions, extensive compatibility tests and a workload-specific jump from 7.55 to 120 session lifecycles per second.

Android Bench 2.0 shows frontier coding agents still fail most multi-day Android tasks

Android Bench 2.0 moves coding-agent evaluation away from small repository fixes toward dependency upgrades, app builds, migrations and other jobs that can take a human engineer days. The results expose a much larger reliability gap than short-task benchmarks—and show that the agent harness can materially change cost and outcome.

Vercel lets coding-agent harnesses use your existing subscriptions without handing tokens to the sandbox

The change separates three things that are often bundled together: the harness, the subscription that pays for it, and the sandbox that executes it. Builders can switch among supported coding agents behind one interface while reusing existing subscription access and reducing credential exposure inside agent runtimes.

Grafana Agent Observability links live agent telemetry to evals and CI regression gates

The observe–test–release loop now has explicit economics: Free and Pro include 30,000 captured generations and 25 million system-initiated AI tokens per month; Pro overages start at $1.50 per 1,000 generations and $2 per million LLM Eval/Guard tokens, while ordinary telemetry is billed separately.

Cloudflare Browser Run can now hard-limit agent sessions to approved hostnames

The useful change is containment rather than another browser-agent feature. Teams can let an agent operate a real browser while constraining its HTTP and HTTPS reach to the site and dependencies the task actually needs, reducing the blast radius of prompt injection, bad tool decisions or untrusted page content.

GPT-6 Astra reaches broad API rollout with 1.05M context — and enterprise computer-use controls

Astra's adoption question is no longer only model capability. Builders can now model its long-context economics and task-level efficiency, while enterprises get a more explicit control plane for computer use. The same release also raises the cyber-safety boundary: OpenAI says Astra is its first model to reach the Preparedness Framework's Critical cybersecurity capability threshold.