Embedding AI Struggles to Understand Physical Measurements, New Study Finds

Photo of author

By Sophia Chen

Researchers have discovered that popular AI systems designed to understand language and meaning—known as embedding models—do not accurately capture basic physical measurements like mass, distance, time, and volume. This finding matters because these physical quantities have clear, objective definitions and relationships, making them a useful test for how well AI models grasp real-world concepts. The study reveals that instead of reflecting true measurement similarities, AI embeddings are often influenced by how similar the words look or sound, which could limit their usefulness in applications requiring precise understanding of quantities.

Key Takeaways

  • Embedding models show only a weak connection between their internal representations and actual physical measurements like mass or time.
  • The way these models measure similarity is often based on superficial string similarities (how words look or sound) rather than true semantic meaning.
  • Attempts to recalibrate or adjust the models to better align with physical measurement concepts did not significantly improve their accuracy.
  • The study highlights peculiar and unexpected patterns in how embeddings handle measurement-related words, suggesting a gap in current AI understanding.

Embedding models are a core technology behind many AI applications, including language translation, search engines, and recommendation systems. These models convert words and phrases into numerical vectors in a high-dimensional space, where the distance between vectors is supposed to reflect semantic similarity—how closely related the meanings of the words are. For example, the words “cat” and “dog” would be closer together than “cat” and “car.”

To test whether these models truly understand physical measurements, the researchers examined how embedding spaces represent quantities like mass (kilograms), distance (meters), time (seconds), and volume (liters). Unlike abstract concepts, these measurements have clear, objective relationships—doubling a length should roughly double the corresponding numeric representation. However, the study found that embedding models often fail to reflect these objective relationships. Instead, the similarity between embeddings frequently corresponded more to how similar the words looked or sounded (string similarity) rather than their actual measured values.

This means that the models might consider “meter” and “metre” (a British spelling variant) as very similar, but not necessarily connect “meter” with “kilogram” in a way that reflects their physical differences and relationships. The researchers also tried recalibrating the similarity measures within the embedding space to better align with physical quantities, but these adjustments produced only minor improvements.

These findings suggest that while embedding models are powerful tools for capturing many types of semantic relationships, they may struggle with concepts grounded in objective, measurable reality. This limitation could affect AI applications in science, engineering, and any domain where precise understanding of quantities is crucial.

Looking ahead, the researchers’ work points to the need for new approaches that explicitly incorporate physical measurement knowledge into AI models. Integrating structured data about units and quantities, or developing specialized embeddings for measurable concepts, could help bridge the gap between linguistic similarity and real-world understanding. As AI continues to expand into technical fields, improving its grasp of fundamental measurements will be an important step toward more reliable and accurate systems.

Based on research published on arXiv by Juri Opitz, Andrianos Michail.

Editor's note

This report is framed around the immediate news and the wider implications for regulators, companies and users following the story.

Article briefing

Researchers have discovered that popular AI systems designed to understand language and meaning—known as embedding models—do not accurately capture basic physical measurements...

Story details

  • Author: Sophia Chen
  • Published: September 18, 2026
  • Category: AI

Key developments

  • Researchers have discovered that popular AI systems designed to understand language and meaning—known as embedding models—do not accurately capture basic physical measurements like mass, distance, time, and volume.
  • Embedding models are a core technology behind many AI applications, including language translation, search engines, and recommendation systems.
  • These models convert words and phrases into numerical vectors in a high-dimensional space, where the distance between vectors is supposed to reflect semantic similarity—how closely related the meanings of the words are.

Why this matters

This finding matters because these physical quantities have clear, objective definitions and relationships, making them a useful test for how well AI models grasp real-world concepts.

Impact and next steps

This limitation could affect AI applications in science, engineering, and any domain where precise understanding of quantities is crucial.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI