As videos grow longer and more complex—capturing hours or even days of activity—getting AI to understand and answer questions about them becomes increasingly challenging. A new research paper introduces an innovative method that helps AI systems remember and track specific objects or people throughout lengthy videos, improving their ability to answer detailed questions about what happened over time.
Key Takeaways
- Researchers developed a system called Grounded Entity Biographies (GEB) that groups observations of the same physical object across different video clips, creating a “biography” of that entity.
- GEB improves AI’s ability to answer questions about long videos by linking events involving the same entity, rather than treating each event in isolation.
- Tests on multiple benchmarks, including videos spanning days, showed GEB outperformed previous memory-based models by significant margins, achieving 72.0% accuracy on the EgoLifeQA dataset.
- The improvements come from combining grounded visual identity tracking with contextual event information; simply adding more descriptive text did not yield the same benefits.
Typical AI models that analyze video content often rely on chronological descriptions or text labels to identify objects and events. However, these methods can struggle when multiple objects share similar descriptions or when the same object appears at different times without explicit links connecting those appearances. For example, a person wearing a red jacket might show up in several scenes across a day, but without a way to confirm it’s the same individual, the AI’s understanding remains fragmented.
The new approach, Grounded Entity Biographies, tackles this problem by visually “grounding” each entity—meaning it ties observations directly to the physical appearance of the object rather than just text labels. The system groups these grounded observations into a continuous memory, effectively creating a timeline or biography for each entity. When a question is asked about the video, the AI can retrieve not only the specific event in question but also the entire history of that entity, allowing it to reason about the entity’s actions and changes over time.
To build these biographies, the researchers’ method analyzes video clips to detect and visually identify entities, linking instances that share the same physical identity across hours or days. This memory is then queried during question answering, combining episodic evidence (what happened when) with the entity’s biography (who or what the entity is, and its past interactions). This dual retrieval enables more accurate and context-aware responses.
In evaluations, the GEB framework was tested on four different benchmarks involving long videos, including the EgoLifeQA dataset, which features day-long recordings. Compared to prior models, GEB consistently improved performance on both multiple-choice and open-ended question answering tasks. The researchers also conducted ablation studies—tests where parts of the system are removed to see their impact—which showed that both the grounded identity linking and the biography retrieval contributed meaningfully to the improved results. Simply adding more descriptive text without identity grounding was not enough to achieve these gains.
This research marks an important step toward AI systems that can better understand and reason about complex, long-duration video content. Potential applications include enhanced video summarization, surveillance analysis, and even helping people search and interact with personal video archives. Future work may explore scaling this approach to even longer videos and more diverse scenarios, as well as integrating it with other AI systems to build richer, more persistent video memories.
Based on research published on arXiv by Hui Ren, Lei Fan, Henry Pao et al..
