TokenCast offers a smarter way to predict how much computational effort large language model agents will use

Photo of author

By Sophia Chen

Large language model (LLM) agents—AI systems that perform tasks by generating and processing text—are becoming increasingly common in applications like virtual assistants, automated research, and coding helpers. However, one tricky challenge has been predicting how many “tokens” (the chunks of text these models process) an agent will consume when completing a task. This matters because token usage directly affects computational cost and speed. A newly published research paper introduces TokenCast, a novel method that forecasts token consumption during an LLM agent’s task execution more accurately and efficiently than before.

Key Takeaways

  • Token consumption by LLM agents can vary widely for the same task, making it difficult to estimate costs ahead of time.
  • TokenCast breaks down the task into segments and learns how much each segment consumes, including the extra cost caused by growing context from earlier steps.
  • It updates its predictions dynamically as the agent runs, without needing extra costly model calls, delivering fast forecasts (about 33 milliseconds on average).
  • Compared to previous methods, TokenCast reduces prediction errors by around 14.5% and can help save over 20% in token usage when managing budgets for task execution.

When an LLM agent works on a task, it often calls on various tools and uses feedback from earlier steps to decide what to do next. Each step adds more information, increasing the “context” the model must consider for subsequent calls. Because this context grows, the number of tokens processed doesn’t just depend on the task itself but also on the sequence of prior actions. This makes predicting total token consumption challenging before or even during execution.

TokenCast addresses this by learning a “composable cost representation” for each part of the task. Think of the task as divided into smaller segments, each with its own token usage and effect on context size. By combining these segments, TokenCast estimates the cumulative token cost, including the extra tokens needed due to the growing context that must be re-read in later steps. This compositional approach allows the prediction to be flexible and updateable as new information appears during the run.

One key advantage is that TokenCast refreshes its forecasts using only the data it already has—without extra calls to the language model, which can be expensive and slow. This makes the system efficient, with an average prediction time of just 32.8 milliseconds on a standard benchmark. The researchers tested TokenCast across multiple task types and agent models, finding consistent improvements over the strongest existing methods.

Beyond more accurate predictions, TokenCast can help manage token budgets more effectively. In scenarios where computational resources or costs need to be controlled, the system enabled agents to use significantly fewer tokens while still completing tasks successfully, compared to fixed-budget approaches.

This research opens the door to more efficient and cost-effective deployment of LLM agents in real-world applications. As language models become integral to various AI-driven tools, being able to forecast and control token consumption will be crucial for scaling these technologies responsibly. The authors have made their code publicly available, inviting further experimentation and development based on their approach.

Based on research published on arXiv by Chaoqian Ouyang, Ling Yue, Libin Zheng et al..

Editor's note

Editors matched this AI update with related coverage to show where it sits in the broader race over models, regulation and product strategy.

Article briefing

When an LLM agent works on a task, it often calls on various tools and uses feedback from earlier steps to decide what to do...

Story details

  • Author: Sophia Chen
  • Published: September 29, 2026
  • Category: AI

Key developments

  • Each step adds more information, increasing the “context” the model must consider for subsequent calls.
  • Because this context grows, the number of tokens processed doesn’t just depend on the task itself but also on the sequence of prior actions.
  • This makes predicting total token consumption challenging before or even during execution.

Why this matters

TokenCast addresses this by learning a “composable cost representation” for each part of the task.

Impact and next steps

As language models become integral to various AI-driven tools, being able to forecast and control token consumption will be crucial for scaling these technologies responsibly.

Background

When an LLM agent works on a task, it often calls on various tools and uses feedback from earlier steps to decide what to do next.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI