Find published dossiers by topic, company, product or technology.

Showing 1–20 of 292 dossiers

Google Cloud opens an agent-ready device farm for mobile testing

Google Cloud’s Developer Device Platform is now in public preview with remote physical-device streaming, parallel emulator testing, smart sharding and an agent skill that can drive multi-step journeys, inspect visual issues and feed fixes back into coding agents. It is billed per active device minute and remains a pre-GA service.

Meta Muse turns a consumer AI assistant into a persistent agent — and its first Mac zero-day tests the containment model

Muse packages persistent autonomous execution, credentials, payments, app access and memory into a mainstream consumer product. A September macOS hotfix now provides an early real-world lesson: agent containment has to protect not only the cloud runtime but also the local control path into the agent.

Google DeepMind is piloting double-blind frontier-model evaluations with confidential computing

The pilot attacks a persistent evaluation trade-off: labs do not want to reveal frontier-model internals, while evaluators do not want benchmark prompts leaking back to the model provider. DeepMind says a Singapore AI Safety Institute pilot kept both sides’ sensitive assets hidden during execution.

Anthropic’s Model Hardware Standard gives AI agents a shared interface for physical devices

MHS is an attempt to make microscopes, liquid handlers, robotic arms and other programmable hardware look like a consistent tool surface to AI agents. It is still a research preview, but the interoperability layer is already being tested with research institutions and hardware vendors.

Appeals court revives the Pentagon’s Anthropic supply-chain blacklist, restoring a Claude procurement barrier

The Anthropic procurement fight changed materially on September 25: a 2–1 federal appeals-court ruling backed the Pentagon’s supply-chain-risk designation. Builders serving defense customers should no longer rely on the August district-court ruling as evidence that the Claude procurement barrier is gone.

Omarchy turns $1.95M in pledged AI credits into an open-source development budget

The interesting part is not another sponsorship total. DHH says Omarchy Quattro is already being built heavily with coding agents, and the token pledges are intended for debugging, security work and a 1,600-plus pull-request backlog. The dollar values are foundation-reported pledged credits, not audited cash spend.

Mixpanel AI can now investigate why a product metric changed and return the result as a working Board

Product teams can launch a root-cause investigation from an Insights report, an alert or Mixpanel Agent instead of manually trying breakdown after breakdown. The result is operationally useful, but it remains an automated statistical diagnosis rather than proof of causation.

AWS Security Agent can now hard-cap autonomous pentest spend and revalidate individual fixes

AWS’s agentic pentesting service can run multiple security tasks in parallel, so billable task-hours may exceed wall-clock test duration. New per-run task-hour limits stop a test gracefully at the ceiling and preserve findings, while targeted revalidation checks specific fixes without rerunning the entire pentest.