What changed
Google says that beginning September 2, 2026, Gemini Notebook consumer accounts will use flexible compute-based usage limits rather than relying only on simpler feature-count quotas. The shared allowance factors the complexity of a prompt, the length of the chat, the number of sources in the notebook and which features are used. Limits refresh every five hours until the account reaches its weekly limit. Gemini Notebook will show usage information and can suggest lower-compute alternatives when an intended output would exceed the remaining allowance. Users who hit the limit can also defer compute-heavy artifacts such as Video Overviews or Slide Decks; Google says those jobs will run automatically later when capacity becomes available, with optional notifications.
Why it matters
This makes AI SaaS usage feel less like 'N prompts per day' and more like spending an invisible compute budget. Two users can issue the same number of prompts and consume different shares of their allowance depending on context length, source count and requested output. That packaging is relevant beyond Gemini Notebook because it shows how providers can expose expensive multimodal and long-context features without publishing a separate per-feature price to consumers. For users, it rewards more deliberate workload planning; for AI SaaS builders, it is another example of product packaging tracking actual computational cost more closely than seats or simple message counts.
Usage becomes workload-sensitive rather than request-count-sensitive
Google says the overall allowance now depends on prompt complexity, chat length, notebook source count and the chosen feature. A short source-grounded question and a long-context multimedia generation can therefore consume materially different amounts of the same budget even though both appear as one user action.
The refresh cycle gets shorter but gains a weekly ceiling
Instead of waiting for a daily reset, users receive refreshed capacity every five hours until the weekly limit is reached. That reduces the chance that one heavy morning session ends the day completely, while still preserving a larger horizon that prevents continuous five-hour resets from creating unlimited use.
Expensive artifacts can be deferred
When a Video Overview or Slide Deck would exceed current capacity, Gemini Notebook can queue the job to run later. Google says deferred outputs generate automatically and users can receive a notification when they are ready. This turns quota exhaustion from a hard error into a scheduling decision for some asynchronous work.
The product can steer users toward cheaper alternatives
Gemini Notebook will track usage and suggest alternative outputs when the first choice would exceed the remaining compute allowance. That makes cost-aware routing part of the interface: the product can trade output type or complexity for immediate availability without exposing a raw token or GPU price.
Higher limits become another upgrade lever
Google’s help documentation says users who hit their limits can wait for refreshes or upgrade for higher AI limits. The company has not published a universal conversion between a prompt and compute units, so customers can see the remaining budget without independently calculating the exact cost of each request in advance.