Azure Document Intelligence v2.0 reaches retirement on August 31, 2026. Microsoft recommends moving workloads to the current v4.0 API; the post-v2 REST surface was redesigned, so teams should verify the actual api-version their SDK or HTTP client sends rather than assuming a package upgrade is enough.
The live DeepSeek changelog and rate card still show distinct V4 Pro service after the previously announced September 14 reroute. That changes cost and model-selection assumptions.
Meta has made the privacy-versus-price trade explicit in its Model API: developers can choose standard pricing or a contributor model ID with steeply discounted inference in exchange for training-data permission. The choice matters for proprietary code, customer data and AI SaaS workloads.
Claude text watermarking is now part of Anthropic’s compliance approach for newly launched models. It does not add tokens or user identifiers, but it is weaker on short, factual, lightly edited and code-heavy outputs, limiting how provenance claims should be used.
The Imagen 4 shutdown is now effective, not merely scheduled. Builders still calling the old model IDs need to migrate to current Gemini image generation, where model names and interaction patterns differ enough to warrant explicit compatibility testing.
The previously reported Stripe–OpenRouter deal is now official. The companies have announced an acquisition agreement, removing the dossier’s main uncertainty; the next questions are closing, product independence, pricing and how deeply Stripe integrates token routing with billing.
Gemini 3.8 Flash keeps 3.7 Flash’s promotional per-token rate and Flash-tier latency, but early independent analysis suggests harder reasoning can increase tokens consumed per task. A separate 3.8 Flash Cyber model is available only through Google’s Fairwind defensive-security program.
Published Updated 5 min read
Inference is where an AI product meets its latency target, reliability budget and monthly bill. Model quality matters, but so do rate limits, caching, batching, regional availability, data terms, observability and the provider behaviour that only appears under production traffic.
BTN tracks important API launches, price changes and serving techniques across hosted and self-managed systems. Coverage connects provider documentation with benchmarks and operating experience so builders can compare more than headline token prices. The useful outcome is knowing when an infrastructure change makes a product newly viable, when migration is worth the work and where apparent savings hide another constraint.
The beat includes routing layers, gateways and compatibility standards when they reduce switching cost or improve control. It also watches changes to retention, abuse monitoring and service terms, because the fastest endpoint is not a safe default if its data handling conflicts with the product being built.