Showing 1–20 of 67 dossiers

AWS Strands Decider 2B turns bounded agent decisions into an open local model

Jev made bounded decision models visible; Strands Decider makes the pattern reproducible inside an agent stack. AWS replaced Qwen3.5-2B's language-generation head with a small scoring head and released the recipe, creating a local alternative for decisions that do not need a full generative model.

Docker turns coding-agent sandboxes into movable cloud compute — and packages their authority as OCI

Docker’s new agent stack combines pay-as-you-go microVM sandboxes with an OCI-based Kit format for declaring what an agent can use. Cloud sessions cost from $0.07 to $1.12 an hour, and Docker says it plans to take the Kit specification toward CNCF neutral governance.

Cloudflare Worker Previews gives every Git branch its own isolated runtime

Branch previews are common for frontend code, but Worker Previews extends the boundary to the runtime itself. Each branch can have independent bindings, state and logs, making parallel human and agent work safer while preserving a production-like execution path.

Claude Code Projects turns one engineering goal into parallel cloud-agent branches

The important change is not simply that Claude can run several agents. Projects now owns decomposition, shared context, branch isolation and progress coordination across full Claude Code sessions, while the trade-offs become usage burn, cloud-only execution and ordinary merge conflicts when parallel work overlaps.

Vercel lets coding-agent harnesses use your existing subscriptions without handing tokens to the sandbox

The change separates three things that are often bundled together: the harness, the subscription that pays for it, and the sandbox that executes it. Builders can switch among supported coding agents behind one interface while reusing existing subscription access and reducing credential exposure inside agent runtimes.

AI agents connect models to tools, memory and multi-step work. That opens useful product possibilities, but it also introduces failure modes that a polished demo can hide: weak recovery, unclear permissions, runaway cost, brittle browser control and uncertain responsibility when an action goes wrong.

This page tracks agent products, protocols, frameworks and research with an eye on real deployment. BTN looks for evidence about reliability, human oversight, security and economics, then translates it into choices a builder can make. The aim is to distinguish durable capability from agent theatre and to keep watching the details that decide whether an agent belongs in production.

The beat also covers the less glamorous work around evaluation and control: permission design, audit trails, approvals, sandboxing and benchmarks that measure completed tasks rather than persuasive transcripts. Agent capability matters most when a team can understand the boundary of what the system may do.