RL Environments Enterprise Huzzle World Talk to founders opens Calendly

Huzzle World: The Frontier of Small Computer-Use Models.

HuzzleWorld ranks number 1 among all models of 2B and smaller.

HuzzleWorld is Huzzle Labs’ family of computer-use models, released at 1B, 2B, 4B and 8B parameters and built on our proprietary architecture, our own RL environments, and training from video. On ScreenSpot-Pro, the hardest public GUI grounding benchmark, they are state of the art for their size, and HuzzleWorld-2B outscores grounding models up to 36 times its size.

HuzzleWorld is #1 among all models of 2B and smaller.

HuzzleWorld-2B Competing 2B models
20 30 40 50 60 70 HuzzleWorld-2B (Zoom In) · 64.2% 64.2 HuzzleWorld-2B (Zoom In) MAI-UI-2B (Zoom In) · 62.8% 62.8 MAI-UI-2B (Zoom In) HuzzleWorld-2B (single pass) · 60.8% 60.8 HuzzleWorld-2B (single pass) UI-Venus-1.5-2B · 57.7% 57.7 UI-Venus-1.5-2B MAI-UI-2B · 57.4% 57.4 MAI-UI-2B GUI-Actor-2VL-2B · 36.7% 36.7 GUI-Actor-2VL-2B UI-TARS-2B · 27.7% 27.7 UI-TARS-2B UGround-V1-2B · 26.6% 26.6 UGround-V1-2B
All 8 leaderboard entries at ≤2BScreenSpot-Pro accuracy %
The result

The best small model for enterprise software.

ScreenSpot-Pro is the leading benchmark for GUI grounding, the task of finding the right thing to click in real professional desktop software. HuzzleWorld-2B scores 64.2%, ahead of every model at 2B or below and ahead of models up to 72B. It was tested on 23 professional applications at full working resolution.

60.8%
Overall, single-pass
Highest of any 2B model
64.2%
With two-pass zoom-in
Optional coarse-to-fine mode
74.8%
Office software
Word, Excel, PowerPoint
70.3%
Text-labelled targets
45.4% on unlabelled icons
HuzzleWorld-2B Qwen2.5-VL-72B-Instruct UI-TARS-72B
OfficeWord · Excel · PowerPoint
74.8
72.6
54.8
DevelopmentVSCode · PyCharm · Android Studio
63.9
53.5
40.8
ScientificMATLAB · Origin · Stata
61.8
59.1
45.7
Operating systemsWindows · macOS · Linux
61.7
49.5
30.1
CreativePhotoshop · Blender · Premiere
54.8
44.9
39.6
CADAutoCAD · SolidWorks · Vivado
51.0
44.4
17.2

HuzzleWorld-2B against two 72B grounding models across all six ScreenSpot-Pro domains. It leads in every one, at 1/36th the parameter count. Bars are percentage accuracy; values are the published single-pass leaderboard results.

Size versus score

HuzzleWorld has moved the frontier for small models.

Every leaderboard entry with a published parameter count, plotted from the largest models on the left to the smallest on the right. The line is the best score achievable at each size. At 2B, HuzzleWorld sits on it.

Leaderboard entry HuzzleWorld-2B Best score at this size or smaller
20 30 40 50 60 70 80 400B 200B 100B 50B 20B 10B 5B 2B AdaZoom-GUI-4B (cond. zoom-in, refine w/ Qwen3.5-397B) · 76.8% · 401B Holo2-235B-A22B (Agentic) · 78.5% · 235B Holo2-235B-A22B · 70.6% · 235B Holo1.5-72B · 63.3% · 72B UI-Venus-72B · 61.9% · 72B GTA1-Qwen2.5VL-72B · 58.4% · 72B Qwen2.5-VL-72B-Instruct · 53.3% · 72B UGround-v1-72B · 34.5% · 72B UI-TARS-72B · 38.1% · 72B Duvo Eye-1 (Holo-3.1-35B-A3B + LoRA) · 72.9% · 35B KV-Ground-8B + Qwen3.5-27B Consistency Router · 80.9% · 35B MAI-UI-32B (MVP) · 77.5% · 32B MAI-UI-32B · 67.9% · 32B MAI-UI-32B (Zoom In) · 73.5% · 32B MVP_Qwen3VL-32B · 74.1% · 32B GTA1-32B · 63.6% · 32B GTA1-Qwen2.5VL-32B · 53.6% · 32B Qwen2.5-VL-32B-Instruct · 48.0% · 32B UI-Venus-1-5-30B-A3B · 69.6% · 30B Holo2-30B-A3B (Agentic) · 75.2% · 30B Holo2-30B-A3B · 66.1% · 30B KV-Ground-GuiOwl1.5-0315-8B-ZoomIn · 80.5% · 8B KV-Ground-GuiOwl1.5-0315-8B · 73.2% · 8B UI-Venus-1-5-8B · 68.4% · 8B Holo2-8B (Agentic) · 71.4% · 8B MAI-UI-8B (Zoom In) · 71.9% · 8B MAI-UI-8B · 65.7% · 8B MVP_Qwen3VL-8B · 65.0% · 8B Holo2-8B · 58.9% · 8B BAMI-7B · 57.7% · 7B UI-AGILE-7B · 47.9% · 7B GTA1-7B · 55.5% · 7B GUI-ARP-7B · 60.8% · 7B Holo1.5-7B · 57.9% · 7B V2P-7B · 52.5% · 7B UI-Venus-7B · 50.8% · 7B TianXi-Action-7B · 51.9% · 7B GUI-Actor-2.5VL-7B · 44.6% · 7B GUI-Actor-2VL-7B · 40.7% · 7B GTA1-Qwen2.5VL-7B · 50.1% · 7B SE-GUI-7B · 47.2% · 7B Qwen2.5-VL-7B-Instruct · 26.8% · 7B UI-TARS-7B · 35.7% · 7B Aguvis-7B · 22.9% · 7B UGround-V1-7B · 31.1% · 7B AdaZoom-GUI-4B (conditional zoom-in) · 70.6% · 4B AdaZoom-GUI-4B · 61.6% · 4B KV-Ground-Qwen3VL-4B-ZoomIn · 70.3% · 4B KV-Ground-GuiOwl1.5-4B-0228-ZoomIn · 76.4% · 4B KV-Ground-Qwen3VL-4B · 63.2% · 4B KV-Ground-GuiOwl1.5-0228-4B · 67.0% · 4B Holo2-4B (Agentic) · 68.6% · 4B Holo2-4B · 57.2% · 4B GUI-AIMA-3B · 59.6% · 3B UI-AGILE-3B · 45.0% · 3B Holo1.5-3B · 51.5% · 3B Di-GUI-3B · 31.6% · 3B ZonUI-3B · 28.7% · 3B GUI-Actor-2.5VL-3B · 42.2% · 3B SE-GUI-3B · 35.9% · 3B UI-Venus-1-5-2B · 57.7% · 2B MAI-UI-2B (Zoom In) · 62.8% · 2B MAI-UI-2B · 57.4% · 2B GUI-Actor-2VL-2B · 36.7% · 2B UI-TARS-2B · 27.7% · 2B UGround-V1-2B · 26.6% · 2B HuzzleWorld-2B · 60.8% HuzzleWorld-2B (Zoom In) · 64.2% HuzzleWorld-2B
Parameters, large to small · log scaleScreenSpot-Pro accuracy %

Moving right means a smaller model. The frontier steps down as size falls, and HuzzleWorld-2B is the point where it ends: nothing smaller scores higher.

In production

Train on your recordings. Run on your infrastructure.

The same properties that make HuzzleWorld small make it practical to deploy inside a company rather than behind someone else’s API.

Your data

Your recordings are the training set

Because the models train from video, we can train directly on the explainer recordings your team already has, and deliver working agents for those workflows. No annotation project first.

Your budget

A fraction of the data and compute

A small model needs far less data, compute and time to train than a large one, while matching its accuracy on the workflows you care about. Shorter path from recording to running agent.

Your infrastructure

Nothing leaves the organization

1B to 8B fits on hardware enterprises already run. Deploy HuzzleWorld inside your own environment, so screen recordings and workflow data never leave it.

Questions

HuzzleWorld, in short.

What is HuzzleWorld?

HuzzleWorld (also written Huzzle World) is Huzzle Labs’ family of computer-use models, released at 1B, 2B, 4B and 8B parameters. Each takes a screenshot and an instruction and returns the exact screen coordinate to act on, which is the grounding step underneath any computer-use agent.

How does HuzzleWorld-2B score on ScreenSpot-Pro?

60.8% single-pass and 64.2% with an optional two-pass zoom-in. It reaches 74.8% on office software and 70.3% on text-labelled targets. Both settings are the highest recorded for a 2B model on the official leaderboard.

Which larger models does it outscore?

Single-pass, HuzzleWorld-2B outscores grounding models up to 36 times its size, including UI-TARS-72B (38.1%), UGround-v1-72B (34.5%), Qwen2.5-VL-72B-Instruct (53.3%), and H Company’s Holo2-8B (58.9%) and Holo1.5-7B (57.9%).

What is ScreenSpot-Pro?

A GUI grounding benchmark built from full-resolution screenshots, up to 3840 × 1080, of 26 real professional applications. A prediction only counts if it lands inside the ground-truth box. There is no partial credit.

Why does model size matter for computer use?

Grounding runs on every step of an agent trajectory, so its cost and latency multiply across a workflow. A 2B model that matches or beats a 72B model changes where computer-use agents can be deployed and what they cost to run.

How is HuzzleWorld trained?

On our proprietary architecture, in our own long-horizon RL environments, and from video. Learning from screen recordings instead of large hand-annotated datasets means the models can be trained directly on a company’s existing explainer recordings.

Can we run it on our own infrastructure?

Yes. At 1B to 8B the models fit on hardware enterprises already run, so HuzzleWorld can be deployed inside your own environment, with no data leaving the organization.

What is GUI grounding?

The step that turns an instruction and a screenshot into a screen coordinate: deciding exactly where to click. It is the part of a computer-use agent that connects a plan to an action, and it runs on every step of a workflow.

Work with us

Deploying computer-use agents, or training your own?
Let’s talk.

Leaderboard data as of 30 July 2026. Ranks are positions among the 72 ScreenSpot-Pro leaderboard entries with a published parameter count and a score of 10 or above. Charts are plotted from 20% upwards; four entries scoring below 20% fall outside the plotted range. Scores are grounding accuracy across the benchmark’s 23 applications and three operating systems; the benchmark measures localisation on a static screen, not end-to-end task completion. Single-pass and zoom-in are recorded as separate leaderboard entries. Figures for HuzzleWorld-1B and HuzzleWorld-4B are our own measurements on the same benchmark, with leaderboard submissions in review.