UI-TARS

ByteDance's open native GUI agent that reads screenshots, not APIs

Agents & Automation Open Source Has API Open Source
Researched · Published · Reviewed
RECATOOLS Score
6.9 / 10
Capability
7
Value for money
8
Ease of use
5
ASEAN readiness
6
API quality
6
Founded
HQ
Users
Launched
Developer

Overview

UI-TARS is ByteDance's open-source, Apache-2.0 licensed GUI agent stack (Agent TARS + UI-TARS Desktop) that operates computers and browsers by reading raw screenshots and issuing clicks, drags and keystrokes — no DOM access or accessibility trees required.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

What you can produce with UI-TARS

  • Native vision-language GUI agent (screenshot in, actions out)
  • Agent TARS: CLI + Web UI for terminal, browser and tool-use tasks
  • UI-TARS Desktop: local computer automation app (Windows/macOS)
  • MCP server integration for real-world tool connectivity
  • Cross-domain coverage: desktop, mobile emulation, browser
  • Open weights, self-hostable, Apache-2.0 licensed
Advertisement

ASEAN Perspective

UI-TARS in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

UI-TARS is one of the few genuinely native computer-use models in open source — it's not a wrapper prompting GPT-4o or Claude to guess at pixel coordinates, it's an end-to-end vision-language model trained specifically to turn a screenshot into the next action. ByteDance's own benchmarks put it ahead of GPT-4o and Claude 3.5 Sonnet on VisualWebBench and ahead of Claude's Computer Use on OSWorld at higher step budgets, with a particular edge on mobile interfaces where rivals struggle.

Real-world reliability is the honest caveat: independent testing on dense, state-dependent interfaces has found accuracy dropping sharply and variance climbing, and ByteDance ships no equivalent of Anthropic's safety-guardrail layer — you're expected to sandbox and supervise it yourself. At 37,900+ GitHub stars it's the most-starred open computer-use project around, free and Apache-2.0 licensed, but documentation is thinner than commercial alternatives and the last tagged release predates its most recent star growth. Good for researchers and technical tinkerers, not yet a drop-in production agent.

Independent AI-assisted assessment by RECATOOLS.

What people say

UI-TARS-desktop is ByteDance's open-source computer-use stack, and it has become one of the fastest-growing open-source AI agent projects of the past two years — from roughly 27,000 stars in a two-week stretch after a viral moment through to 37,900+ stars and 3,800 forks by mid-2026. It ships under Apache-2.0 and packages two related projects: Agent TARS, a CLI-and-Web-UI agent for terminal, browser and general tool-use tasks, and UI-TARS Desktop, a native app that controls the local machine (screenshots in, mouse/keyboard actions out).

The technical distinction that reviewers keep coming back to: UI-TARS is trained end-to-end as a vision-language model that takes a raw screenshot as its only input and outputs clicks, drags, typed text and scroll actions — no DOM parsing, no accessibility-tree access, no per-app integration required. That makes it generalizable across desktop apps, mobile emulators and browsers in a way DOM-dependent browser agents aren't. ByteDance's published benchmarks show it leading VisualWebBench (82.8%, ahead of GPT-4o and Claude 3.5 Sonnet) and outperforming Claude's Computer Use on OSWorld at higher step budgets (24.6 vs. 22.0) — and ByteDance specifically credits it with holding up better than Claude Computer Use on mobile-scenario tasks, where DOM-less screenshot models have a structural edge.

Independent evaluation is less flattering on harder cases: accuracy on state-dependent and visually dense interfaces has been measured falling to around 38% in some third-party testing, with high variance — the gap between 'wins the benchmark' and 'reliable in a messy real app' is real. ByteDance's own technical reports (UI-TARS-2, published via arXiv) acknowledge this and describe multi-turn reinforcement learning as the fix path, an ongoing research effort rather than a solved problem.

Practically: the desktop app and CLI are both free, and compute cost depends entirely on which model backend you point it at — self-hosted UI-TARS weights, a hosted UI-TARS endpoint, or another vision-capable model via the SDK. Documentation is noticeably thinner than Anthropic's computer-use materials, there's no mature plugin/extension ecosystem yet, and safety guardrails are left to the operator rather than built in. The latest tagged release (v0.3.0) dates to November 2025, though the project has continued shipping via commits and companion repos since.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including UI-TARS's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with UI-TARS unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to UI-TARS directly →

Spotted something out of date? Suggest an update →

Advertisement