Coding benchmark
SWE-bench Multimodal
Built to test whether coding agents generalize beyond Python to visual front end, diagramming, mapping and data visualization codebases. The repo notes that Multimodal v2 is now fully open source.
SWE-bench team (Princeton University and collaborators)
At a glance
- What the agent must do
- The agent must fix bugs in user facing JavaScript libraries where the issue includes visual elements such as screenshots.
- Size
- 617 tasks in the paper; 480 tasks open for local evaluation in Multimodal v2
- Scoring
- % of issues resolved (tests pass)
- How to run it
- Docker harness via the SWE-bench repo using the multimodal dataset alias
Official sources
Other coding benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.