ESPO: A Smarter Way to Fine-Tune AI Prompts Boosts Accuracy and Cuts Length

Photo of author

By Sophia Chen

Researchers have developed a new method called ESPO that improves how AI language models are guided to answer questions and perform tasks. When interacting with AI, “prompts” are the instructions or examples given to the model to help it generate the right response. However, existing techniques for optimizing these prompts often make them longer and more complicated without improving accuracy. ESPO tackles this problem by diagnosing errors more effectively, generating diverse prompt variations, and selecting the best candidates with a stability check. This approach leads to shorter, more accurate prompts that work better across different AI models.

Key Takeaways

  • ESPO improves average accuracy on seven public natural language benchmarks by nearly 4 percentage points compared to previous state-of-the-art methods.
  • It produces prompts that are about 47% shorter, reducing complexity and speeding up AI responses.
  • The method works consistently well across multiple large AI models, including Gemma 3, Mistral, Qwen3, and Claude Haiku.
  • Incorporating diverse candidate prompts without a stability selection step can actually decrease performance, highlighting the importance of ESPO’s three-phase approach.

At the heart of ESPO is a three-step process designed to fix key shortcomings in earlier prompt optimization methods like GEPA. First, the “Diagnose” phase examines where the AI model makes mistakes and groups these errors into structured patterns. Instead of blindly expanding prompts, this helps focus improvements on known weaknesses. Next, the “Propose” phase generates new prompt candidates using four different strategies, each bringing a unique perspective to the search for better instructions. Finally, the “Select” phase uses a statistical technique called bootstrap stability selection to reliably pick prompt variants that consistently improve performance across different data splits, reducing the risk of overfitting or random success.

Prompt optimization involves adjusting the input instructions to AI models so they produce more accurate outputs. Previous evolutionary approaches tended to add more rules and conditions with every iteration, leading to “prompt bloat” — long, unwieldy prompts that do not necessarily help the model perform better. ESPO’s innovation lies in carefully structuring the error analysis and diversifying candidate generation while ensuring robust selection. This balanced approach results in prompts that are not only more effective but also more concise, making AI systems faster and easier to use.

The researchers tested ESPO on seven well-known natural language processing benchmarks, including tasks like answering questions from tweets, solving math problems, and reasoning over multiple documents. Across the board, ESPO outperformed GEPA, the previous leading prompt optimizer, by producing shorter prompts with higher accuracy. Moreover, when applied to different AI models of varying sizes and architectures, ESPO consistently achieved the best results, sometimes improving accuracy dramatically — for example, on Qwen3’s math problems, accuracy jumped from 76.4% to 91.4%.

Looking ahead, ESPO’s framework offers a promising path to making AI assistants smarter and more efficient by refining how we communicate with them. Shorter, more accurate prompts mean quicker responses and less computational load, which could benefit applications ranging from customer service chatbots to automated research assistants. Future work may explore extending ESPO to other AI tasks beyond natural language or integrating it directly into model training to further enhance AI capabilities.

Based on research published on arXiv by Lihao Liu, Peng Tang, Kunwar Yashraj Singh et al..

Editor's note

This article focuses on the confirmed update first, then points readers to the competitive and policy context that shapes the beat.

Article briefing

Researchers have developed a new method called ESPO that improves how AI language models are guided to answer questions and perform...

Story details

  • Author: Sophia Chen
  • Published: September 4, 2026
  • Category: AI

Key developments

  • Researchers have developed a new method called ESPO that improves how AI language models are guided to answer questions and perform tasks.
  • When interacting with AI, “prompts” are the instructions or examples given to the model to help it generate the right response.
  • However, existing techniques for optimizing these prompts often make them longer and more complicated without improving accuracy.

Why this matters

Shorter, more accurate prompts mean quicker responses and less computational load, which could benefit applications ranging from customer service chatbots to automated research assistants.

Impact and next steps

Next, the “Propose” phase generates new prompt candidates using four different strategies, each bringing a unique perspective to the search for better instructions.

Background

At the heart of ESPO is a three-step process designed to fix key shortcomings in earlier prompt optimization methods like GEPA.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI