Find published dossiers by topic, company, product or technology.

Showing 181–200 of 229 dossiers

Cognition’s Fusion uses a frontier lead and cheaper sidekick to cut coding-agent task cost

Fusion is interesting less as another routing feature than as a different agent-cost architecture: two persistent model contexts divide planning, review and execution instead of making one expensive model handle every token. The practical question for builders is shifting from token price to cost per completed task.

Reddit opens Max campaign creation to third-party Ads API clients

The change makes Reddit's more automated campaign type usable by agencies, ad-tech platforms and internal campaign systems instead of only through first-party buying surfaces. It expands automation reach, but third-party builders inherit Max's creative and optimization assumptions rather than gaining a new manual campaign type.

Google DeepMind is piloting double-blind frontier-model evaluations with confidential computing

The pilot attacks a persistent evaluation trade-off: labs do not want to reveal frontier-model internals, while evaluators do not want benchmark prompts leaking back to the model provider. DeepMind says a Singapore AI Safety Institute pilot kept both sides’ sensitive assets hidden during execution.

A preregistered field experiment finds Google’s AI search reduces publisher clicks

The study moves the AI-search traffic debate beyond observational correlations: participants were randomly assigned to current Google Search, a version with AI features hidden, or AI Mode-only search during ordinary browsing. It is still a preprint and does not establish effects for every query or publisher.

GitHub’s OAuth apps get short-lived tokens, multiple callbacks — and a wildcard setting worth auditing

GitHub OAuth apps can now use eight-hour access tokens with rotating refresh tokens, register up to 10 callback URLs, and explicitly control wildcard callback matching. New apps default to expiring tokens, while existing single-callback apps should review a legacy wildcard setting GitHub has now made visible.

Claude Code Projects turns one engineering goal into parallel cloud-agent branches

The important change is not simply that Claude can run several agents. Projects now owns decomposition, shared context, branch isolation and progress coordination across full Claude Code sessions, while the trade-offs become usage burn, cloud-only execution and ordinary merge conflicts when parallel work overlaps.

Android Bench 2.0 shows frontier coding agents still fail most multi-day Android tasks

Android Bench 2.0 moves coding-agent evaluation away from small repository fixes toward dependency upgrades, app builds, migrations and other jobs that can take a human engineer days. The results expose a much larger reliability gap than short-task benchmarks—and show that the agent harness can materially change cost and outcome.

Jev becomes Vercel AI Gateway’s fastest-adopted model in its first 24 hours

Jev’s launch claims were interesting; Vercel’s usage data is more useful. Nearly 13% of paid AI Gateway teams tried the typed decision model in its first day, while Jev also rose to a material share of gateway requests. That does not establish retention or production success, but it is unusually fast developer uptake for a model designed to make bounded software decisions rather than generate prose.