Showing 41–60 of 67 dossiers

Anthropic’s Model Hardware Standard gives AI agents a shared interface for physical devices

MHS is an attempt to make microscopes, liquid handlers, robotic arms and other programmable hardware look like a consistent tool surface to AI agents. It is still a research preview, but the interoperability layer is already being tested with research institutions and hardware vendors.

TRACE gives AI agents a portable, hardware-attested runtime evidence format

TRACE targets a gap between audit promises and what an AI agent actually did at runtime. Its v0.2 developer preview can bind model, policy, data and tool-use claims to confidential-computing attestation, but it is still pre-ratification and explicitly not ready to treat as a production compliance guarantee.

OpenAI’s Codex client exposes a GenUI layer for refreshable interfaces inside conversations

RuntimeWire found a generic `genui` message path, a server-directed widget refresh endpoint and 467 versioned Learning Block manifests inside OpenAI’s Codex desktop client. The material development is not another visualization feature: it is evidence of a reusable interface layer beneath conversational answers, with important limits around what is actually public or enabled.

AWS Security Agent can now hard-cap autonomous pentest spend and revalidate individual fixes

AWS’s agentic pentesting service can run multiple security tasks in parallel, so billable task-hours may exceed wall-clock test duration. New per-run task-hour limits stop a test gracefully at the ceiling and preserve findings, while targeted revalidation checks specific fixes without rerunning the entire pentest.

Supabase makes MCP access centrally governed through enterprise SSO

Supabase has implemented MCP Enterprise-Managed Authorization using identity-provider assertions, short-lived tokens and existing Supabase role boundaries. It gives organizations a central on/off switch for approved AI clients while keeping access scoped to the individual employee rather than sharing a powerful organization token.

incident.io has made autonomous incident investigations generally available

Investigations has crossed from preview into production and incident.io now reports a large latency improvement in its own measured workflow. The agent continuously reassesses evidence and can hand remediation to coding agents, but the new speed and accuracy figures remain vendor-produced rather than independent.

Google Cloud opens an agent-ready device farm for mobile testing

Google Cloud’s Developer Device Platform is now in public preview with remote physical-device streaming, parallel emulator testing, smart sharding and an agent skill that can drive multi-step journeys, inspect visual issues and feed fixes back into coding agents. It is billed per active device minute and remains a pre-GA service.

GitLab 19.3 turns plain-English process knowledge into runnable agentic flows

Custom Flows became generally available in GitLab 19.2; 19.3 adds the missing authoring layer. Flow Creator reads current Flow Registry docs, applies known failure rules and generates a runnable flow from plain English. Builders still need to review, register and govern the automation rather than treating generated YAML as trusted infrastructure.

Grafana Agent Observability links live agent telemetry to evals and CI regression gates

The observe–test–release loop now has explicit economics: Free and Pro include 30,000 captured generations and 25 million system-initiated AI tokens per month; Pro overages start at $1.50 per 1,000 generations and $2 per million LLM Eval/Guard tokens, while ordinary telemetry is billed separately.

AI agents connect models to tools, memory and multi-step work. That opens useful product possibilities, but it also introduces failure modes that a polished demo can hide: weak recovery, unclear permissions, runaway cost, brittle browser control and uncertain responsibility when an action goes wrong.

This page tracks agent products, protocols, frameworks and research with an eye on real deployment. BTN looks for evidence about reliability, human oversight, security and economics, then translates it into choices a builder can make. The aim is to distinguish durable capability from agent theatre and to keep watching the details that decide whether an agent belongs in production.

The beat also covers the less glamorous work around evaluation and control: permission design, audit trails, approvals, sandboxing and benchmarks that measure completed tasks rather than persuasive transcripts. Agent capability matters most when a team can understand the boundary of what the system may do.