A benchmark for how language models handle real insurance work. It spans two core workflows — underwriting and claims & coverage — built from document-grounded cases that each resolve to a single verifiable answer. Models are evaluated pass@1 and scored against the recorded outcome, not the wording of the response.
Where AI makes contact with the real world.
We build RL environments for frontier AI labs, and the evaluations and custom models that put AI to work inside the enterprise.
| #1 | HuzzleWorld-2B | 64.2% | |
| #2 | MAI-UI-2B | 62.8% | |
| #3 | UI-Venus-1.5-2B | 57.7% | |
| #4 | GUI-Actor-2VL-2B | 36.7% | |
| #5 | UI-TARS-2B | 27.7% | |
| #6 | UGround-V1-2B | 26.6% | |
| #7 | ShowUI | 7.7% |
Huzzle Labs is behind Huzzle World, our family of computer-use models at 1B to 8B, built on a proprietary architecture, our own RL environments and training from video. HuzzleWorld-2B is #1 on ScreenSpot-Pro among all models of 2B and smaller, and single-pass it outscores grounding models up to 36 times its size.
Huzzle Labs is one of the leading AI platforms providing post-training data and infra for AI labs.
Emre Guven
Meta Superintelligence Labs
Training frontier models, or putting them to work in production? Let’s talk.
Benchmarks
We are building benchmarks that test real professional work, starting with InsureBench, a proposed benchmark that measures how models perform on claims management and underwriting workflows.
Backed by 10X Founders·Angel Invest·Emerge·a16z Scout Fund·Thomas Wolf Hugging Face·Bernd Heinemann Allianz·Yaser Khalighi Stanford