Your recordings are the training set
Because the models train from video, we can train directly on the explainer recordings your team already has, and deliver working agents for those workflows. No annotation project first.
HuzzleWorld ranks number 1 among all models of 2B and smaller.
HuzzleWorld is Huzzle Labs’ family of computer-use models, released at 1B, 2B, 4B and 8B parameters and built on our proprietary architecture, our own RL environments, and training from video. On ScreenSpot-Pro, the hardest public GUI grounding benchmark, they are state of the art for their size, and HuzzleWorld-2B outscores grounding models up to 36 times its size.
ScreenSpot-Pro is the leading benchmark for GUI grounding, the task of finding the right thing to click in real professional desktop software. HuzzleWorld-2B scores 64.2%, ahead of every model at 2B or below and ahead of models up to 72B. It was tested on 23 professional applications at full working resolution.
HuzzleWorld-2B against two 72B grounding models across all six ScreenSpot-Pro domains. It leads in every one, at 1/36th the parameter count. Bars are percentage accuracy; values are the published single-pass leaderboard results.
Every leaderboard entry with a published parameter count, plotted from the largest models on the left to the smallest on the right. The line is the best score achievable at each size. At 2B, HuzzleWorld sits on it.
Moving right means a smaller model. The frontier steps down as size falls, and HuzzleWorld-2B is the point where it ends: nothing smaller scores higher.
The same properties that make HuzzleWorld small make it practical to deploy inside a company rather than behind someone else’s API.
Because the models train from video, we can train directly on the explainer recordings your team already has, and deliver working agents for those workflows. No annotation project first.
A small model needs far less data, compute and time to train than a large one, while matching its accuracy on the workflows you care about. Shorter path from recording to running agent.
1B to 8B fits on hardware enterprises already run. Deploy HuzzleWorld inside your own environment, so screen recordings and workflow data never leave it.
HuzzleWorld (also written Huzzle World) is Huzzle Labs’ family of computer-use models, released at 1B, 2B, 4B and 8B parameters. Each takes a screenshot and an instruction and returns the exact screen coordinate to act on, which is the grounding step underneath any computer-use agent.
60.8% single-pass and 64.2% with an optional two-pass zoom-in. It reaches 74.8% on office software and 70.3% on text-labelled targets. Both settings are the highest recorded for a 2B model on the official leaderboard.
Single-pass, HuzzleWorld-2B outscores grounding models up to 36 times its size, including UI-TARS-72B (38.1%), UGround-v1-72B (34.5%), Qwen2.5-VL-72B-Instruct (53.3%), and H Company’s Holo2-8B (58.9%) and Holo1.5-7B (57.9%).
A GUI grounding benchmark built from full-resolution screenshots, up to 3840 × 1080, of 26 real professional applications. A prediction only counts if it lands inside the ground-truth box. There is no partial credit.
Grounding runs on every step of an agent trajectory, so its cost and latency multiply across a workflow. A 2B model that matches or beats a 72B model changes where computer-use agents can be deployed and what they cost to run.
On our proprietary architecture, in our own long-horizon RL environments, and from video. Learning from screen recordings instead of large hand-annotated datasets means the models can be trained directly on a company’s existing explainer recordings.
Yes. At 1B to 8B the models fit on hardware enterprises already run, so HuzzleWorld can be deployed inside your own environment, with no data leaving the organization.
The step that turns an instruction and a screenshot into a screen coordinate: deciding exactly where to click. It is the part of a computer-use agent that connects a plan to an action, and it runs on every step of a workflow.
Deploying computer-use agents, or training your own?
Let’s talk.
Leaderboard data as of 30 July 2026. Ranks are positions among the 72 ScreenSpot-Pro leaderboard entries with a published parameter count and a score of 10 or above. Charts are plotted from 20% upwards; four entries scoring below 20% fall outside the plotted range. Scores are grounding accuracy across the benchmark’s 23 applications and three operating systems; the benchmark measures localisation on a static screen, not end-to-end task completion. Single-pass and zoom-in are recorded as separate leaderboard entries. Figures for HuzzleWorld-1B and HuzzleWorld-4B are our own measurements on the same benchmark, with leaderboard submissions in review.