Coding benchmark

SWE-bench Multimodal

Built to test whether coding agents generalize beyond Python to visual front end, diagramming, mapping and data visualization codebases. The repo notes that Multimodal v2 is now fully open source.

SWE-bench team (Princeton University and collaborators)

At a glance

What the agent must do
The agent must fix bugs in user facing JavaScript libraries where the issue includes visual elements such as screenshots.
Size
617 tasks in the paper; 480 tasks open for local evaluation in Multimodal v2
Scoring
% of issues resolved (tests pass)
How to run it
Docker harness via the SWE-bench repo using the multimodal dataset alias

Official sources

Other coding benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.