Skip to content
Back to Guavy Wire
Crypto

Meta Introduces GAMUT Benchmark to Evaluate Factual Completeness in AI Responses

Instruments
MEW
Share

Meta has introduced a new benchmark called GAMUT to measure the factual completeness of AI responses. The Grounded Assessment of Multimodal Factuality (GAMUT) benchmark aims to evaluate whether an AI's answer includes all the relevant facts, not just if it is accurate or fluent.

The researchers used a two-level meta-rubric framework that organizes required content hierarchically and translates them into yes-or-no questions. This allows language models to grade their performance reliably. The benchmark contains 1,813 questions rooted in real wearable imagery across ten distinct domains.

Meta evaluated 14 different AI models against the GAMUT benchmark, with the top performer being Gemini 3.1 Pro, which scored 58.7%. The results demonstrated strong discriminative ability, meaning it could separate better performers from worse ones.

More on Crypto

Disclaimer: Guavy is a data and market intelligence provider, not an investment advisor. The information, signals, and market analysis provided by the Guavy API and related services are for informational purposes only and are not intended as financial advice, investment recommendations, or an endorsement of any particular trading strategy. Trading in volatile markets, including cryptocurrency, carries significant risk and may not be suitable for all investors. Past performance is not indicative of future results. Users should consult with a qualified financial professional before making any investment decisions. Guavy makes no guarantee of trading profits or financial returns.

Real-time market sentiment intelligence for apps, funds & agents

Location

729 55 Ave SW
Calgary AB T2V 0G4
Canada

© 2026 Guavy Inc