Showing 21–33 of 33 dossiers

Google DeepMind is piloting double-blind frontier-model evaluations with confidential computing

The pilot attacks a persistent evaluation trade-off: labs do not want to reveal frontier-model internals, while evaluators do not want benchmark prompts leaking back to the model provider. DeepMind says a Singapore AI Safety Institute pilot kept both sides’ sensitive assets hidden during execution.

NVIDIA formally agrees to acquire Hugging Face for $12.93 billion — and promises to keep it open across rival hardware

The previously reported NVIDIA–Hugging Face deal is now a definitive agreement rather than an unconfirmed report. The most important new detail for builders is not only the price: NVIDIA has put multi-model and multi-silicon openness into its public and regulatory framing, while the acquisition still faces closing conditions and regulatory approval.

GLM-5.3-Flash turns the anonymous Ox Alpha trial into an open-weight multimodal coding model

GLM-5.3-Flash combines open weights, multimodal coding/agent capability and an 18B-active sparse architecture with a large anonymous pre-launch trial. Z.ai has already issued a chat-template correction for early downloads, showing that day-one self-hosted deployments need artifact-level validation as well as model benchmarking.

Gemini 3.8 Flash raises agent capability at the same token price — but may use more tokens per task

Gemini 3.8 Flash keeps 3.7 Flash’s promotional per-token rate and Flash-tier latency, but early independent analysis suggests harder reasoning can increase tokens consumed per task. A separate 3.8 Flash Cyber model is available only through Google’s Fairwind defensive-security program.

AI models are the engines underneath many new products, but a model launch rarely tells you enough to choose one. This page follows frontier and specialist models, context windows, multimodal capability, evaluation results, pricing and the practical constraints that appear once a model leaves the demo.

BTN compares primary model cards and documentation with credible independent testing. The focus is on decisions: whether a release changes what can be built, whether a benchmark reflects real work, what the serving costs imply, and which limitations still matter. The result is a running view of model progress without treating every leaderboard movement as a breakthrough.

Expect coverage to connect model behaviour with the surrounding product decision. That includes fine-tuning and retrieval options, safety controls, regional access and the pace at which preview features become dependable APIs. Older models stay relevant when lower price or easier hosting makes them the sensible production choice.