MidTool Boosts AI’s Ability to Use Digital Tools More Effectively Through Specialized Training

Photo of author

By Sophia Chen

Researchers have developed a new method to improve how large language models (LLMs) like AI assistants learn to use digital tools, such as software APIs and workflow systems. This mid-training approach, called MidTool, aims to teach these models not just to understand language, but also to recognize what tools can do, how to apply them in context, and how to recover when information is missing. The result is AI that can perform more complex, tool-based tasks with greater accuracy and reliability.

Key Takeaways

  • MidTool introduces a specialized mid-training phase that focuses on teaching AI models general tool use, rather than relying solely on later fine-tuning steps.
  • The method combines large-scale data from web pages, PDFs, and code with synthesized supervision derived from real-world tool APIs and workflows.
  • Models trained with MidTool showed consistent improvements on multiple benchmarks measuring tool use capabilities, outperforming baseline models.
  • This approach helps AI understand tool affordances (what tools can do), how to compose sequences of tool calls, and how to handle incomplete or ambiguous information.

Large language models are typically trained in stages, starting with broad learning on vast text datasets, followed by fine-tuning on specific tasks. Recently, researchers have recognized an important intermediate step called “mid-training,” which can refine certain skills before final fine-tuning. While mid-training has been used to boost reasoning and scientific problem-solving abilities, its potential to improve general tool use—how AI interacts with software tools and APIs—has been less explored until now.

The MidTool approach builds a comprehensive training dataset called MidTool-Mix by gathering diverse information from the internet, PDF documents, and code repositories. It also synthesizes supervision signals based on interactions with real-world tool APIs, multi-step compositional planning (MCP) skills, and workflows grounded in documents. This combination allows the model to learn not just from static text, but from examples that mimic actual tool usage scenarios.

In practical terms, “tool affordances” refer to what actions a tool allows—like a calculator performing arithmetic or a database API retrieving records. Teaching the model to recognize these affordances helps it decide which tool to use for a given problem. Additionally, MidTool trains models to “compose tool call workflows,” meaning the AI learns how to chain multiple tool operations together step-by-step to solve complex tasks. The training also emphasizes recovering from incomplete information, which is crucial in real-world situations where data might be missing or ambiguous.

The researchers applied MidTool mid-training to two versions of the Qwen3 language model and then refined those models further using supervised fine-tuning and reinforcement learning. When tested on benchmarks such as BFCL, tau2-Bench, and MCP Universe—each designed to measure different aspects of tool use—the MidTool-trained models consistently outperformed baseline counterparts. This demonstrates that dedicating a specific mid-training phase to general tool use can significantly enhance AI capabilities.

Looking ahead, this research suggests a promising direction for developing AI systems that can better assist with practical, tool-based tasks across domains like software engineering, data analysis, and automation. By incorporating mid-training focused on tool use, future models might become more reliable collaborators in complex workflows that require interacting with multiple software tools. While further work is needed to expand the range of tools and real-world scenarios, MidTool provides a valuable framework for advancing AI’s agency in digital environments.

Based on research published on arXiv by Fengqing Jiang, Yite Wang, Boyi Liu et al..

Editor's note

Editors matched this AI update with related coverage to show where it sits in the broader race over models, regulation and product strategy.

Article briefing

Researchers have developed a new method to improve how large language models (LLMs) like AI assistants learn to use digital tools, such as software APIs and workflow...

Story details

  • Author: Sophia Chen
  • Published: August 21, 2026
  • Category: AI

Key developments

  • Researchers have developed a new method to improve how large language models (LLMs) like AI assistants learn to use digital tools, such as software APIs and workflow systems.
  • The result is AI that can perform more complex, tool-based tasks with greater accuracy and reliability.
  • Large language models are typically trained in stages, starting with broad learning on vast text datasets, followed by fine-tuning on specific tasks.

Impact and next steps

The training also emphasizes recovering from incomplete information, which is crucial in real-world situations where data might be missing or ambiguous.

Background

Recently, researchers have recognized an important intermediate step called "mid-training," which can refine certain skills before final fine-tuning.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI