Hugging Face · Agent Skill

huggingface-vision-trainer

Trains and fine-tunes vision models for object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models — MobileNetV3, MobileViT, ResNet, ViT/DINOv3 — plus any Transformers classifier), and SAM/SAM2 segmentation using Hugging Face Transformers on Hugging Face Jobs cloud…

Published by Hugging FaceApache-2.0593 lines in SKILL.md
Source
huggingface/skills/skills/huggingface-vision-trainer
Repository owner
huggingface

Install

Read the SKILL.md and any scripts before installing: a skill runs with your agent's permissions. This copies just this skill into Claude Code's user skills folder; other agents read skills from their own folder.

Claude Code (user-level)

git clone --depth 1 --filter=blob:none --sparse https://github.com/huggingface/skills.git /tmp/skills
cd /tmp/skills && git sparse-checkout set "skills/huggingface-vision-trainer"
mkdir -p ~/.claude/skills && cp -r "skills/huggingface-vision-trainer" ~/.claude/skills/huggingface-vision-trainer

Skills folder docs:

Works with

Agent Skills is an open format, so this skill loads in any harness that supports it, including Claude Code, Codex, Gemini CLI, OpenCode, Cursor, GitHub Copilot, goose, OpenHands. Some skills are written for one product and say so in their description.

More skills from Hugging Face

Reviewed Oct 5, 2026. Name and description are the skill's own frontmatter.