Portable agent skills for training, evaluating, and running inference on NVIDIA TAO models. Works with Claude Code, Codex, Gemini CLI, or any coding agent that speaks the Agent Skills open standard.
Overview
Portable agent skills for training, evaluating, and running inference on NVIDIA TAO models. Works with Claude Code, Codex, Gemini CLI, or any coding agent that speaks the Agent Skills open standard.
README
NVIDIA TAO Skill Bank
Portable agent skills for training, evaluating, and running inference on NVIDIA TAO models. Works with Claude Code, Codex, Gemini CLI, or any coding agent that speaks the Agent Skills open standard. Most local Docker model/data actions need only Docker plus NVIDIA Container Toolkit; application workflows may declare small host-Python requirements for deterministic state and analysis adapters. Advanced features—job tracking, multi-node, S3 I/O—are built in, not bolted on: platform skills implement a four-verb execution contract over their native CLI with no nvidia-tao-sdk. See Execution: no SDK required.
Branching strategy
main is our development trunk, not a stable branch — it moves
continuously and can reference pre-release container builds. Each release is
cut onto a release/X.Y.Z branch, QA’d there, and published as a matching
Git tag on the Releases page.
Install from the latest published tag — that is what every command below
pins to — and treat both main and an unpublished release/* branch as
work in progress. Contributors: see
CONTRIBUTING.md for how
changes flow from main into a release.
Known issue — internal image gating. Skills on
main(and on arelease/*branch before its tag is published) may pin release-candidate container images undernvcr.io/nvstaging/*. Those are internal builds: an external NGC account cannot pull them, however valid itsNGC_KEY. Publishing a release rewrites those pins to publicnvcr.io/nvidia/taoimages, so no published tag pins annvstagingimage. We are working on a way to host dev and nightly images publicly; until then, install from the latest published tag.
Install
The skill bank works with both Claude Code and Codex. Pick the runtime you use.
The commands below install the latest release build, 7.1.0. Every install
path pins the marketplace to that tag, so a fresh install gets the same
validated build every time instead of whatever main happens to hold. Newer
tags appear on the Releases page;
substitute the tag you want, or use @main if you specifically need unreleased
work.
Claude Code
In a Claude Code session, add the marketplace at the release tag and install the plugin:
/plugin marketplace add NVIDIA-TAO/[email protected]
/plugin install tao-skills@tao-skill-bank
That’s it — no git clone, no pip install. The TAO Skill Bank plugin bundles all skills (every model, data, platform, and application). The plugin’s SessionStart hook loads the AGENTS.md identity at the start of every session.
Codex
Codex setup has two independent pieces — the plugin (which surfaces the skills to Codex) and AGENTS.md (which loads the agent identity). You need both for parity with Claude Code.
One command (recommended)
curl -fsSL https://raw.githubusercontent.com/NVIDIA-TAO/tao-skill-bank/main/scripts/install-codex-agents.sh | bash
The installer is fetched from main on purpose — it is the copy that knows
which tag is current, and it registers the marketplace at that release tag. A
copy fetched from a release tag only knows about refs that existed when the tag
was cut.
…or, if you’ve already cloned or extracted the repo from a zip, run
scripts/install-codex-agents.sh from that directory. The script registers the
marketplace at the latest release tag (7.1.0), installs the TAO Skill Bank
plugin, and copies AGENTS.md to ~/.codex/AGENTS.md so the TAO identity loads
in every Codex session. It’s idempotent and backs up any existing
~/.codex/AGENTS.md before overwriting. Override the source with
TAO_SKILL_BANK_MARKETPLACE=… and TAO_SKILL_BANK_REF=… to use a fork, a
different tag or branch (TAO_SKILL_BANK_REF=main for unreleased work), or a
local absolute path:
cd /absolute/path/to/tao-skills-external
TAO_SKILL_BANK_MARKETPLACE=/absolute/path/to/tao-skills-external \
scripts/install-codex-agents.sh
Manual steps
If you’d rather drive each step yourself:
1. Install the plugin. Either use the VS Code Codex extension’s plugin UI (select TAO Skill Bank), or from the CLI:
codex plugin marketplace add NVIDIA-TAO/[email protected]
codex plugin add tao-skill-bank@tao-local-plugins
This installs the bundle to ~/.codex/plugins/cache/tao-local-plugins/tao-skill-bank// (the tao-local-plugins segment comes from the name field in .agents/plugins/marketplace.json).
For a local zip or clone, use the absolute path instead of the Git URL:
codex plugin marketplace add /absolute/path/to/tao-skills-external
codex plugin add tao-skill-bank@tao-local-plugins
2. Load the agent identity (AGENTS.md). The plugin install does not auto-load AGENTS.md — Codex’s AGENTS.md discovery walks down from the project root, not into the plugin cache (see openai/codex#16430 for why plugin-bundled SessionStart hooks don’t fix this yet). Pick one:
- Per-project:
git clonethis repo and launchcodexfrom inside the clone. Codex auto-loadsAGENTS.mdfrom the project root per the agents.md cross-runtime spec. - Globally (one-time copy):
cp ~/.codex/plugins/cache/tao-local-plugins/tao-skill-bank//AGENTS.md ~/.codex/AGENTS.md. The identity then loads in every Codex session, anywhere.
Once Codex starts honoring plugin-bundled hooks, the identity will install automatically alongside the plugin — until then, this manual step is needed.
Credentials
The skill bank reads credentials from the session environment — export what you need in your shell before launching, and the session inherits them:
export NGC_KEY=... # nvcr.io image pulls
export HF_TOKEN=... # gated HuggingFace models
The vars each skill looks for (export only the ones your workflow needs):
| Var | Used for |
|---|---|
NGC_KEY |
nvcr.io image pulls — required by almost everything |
HF_TOKEN |
gated HuggingFace models / push_to_hub |
BREV_API_TOKEN |
tao-run-on-brev (optional — brev login also works) |
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, S3_BUCKET_NAME, AWS_ENDPOINT_URL, AWS_DEFAULT_REGION |
S3 / object-storage I/O via tao-data-io (legacy ACCESS_KEY/SECRET_KEY still accepted) |
WANDB_API_KEY, WANDB_PROJECT |
WandB experiment logging (AutoML / HF fine-tune) |
Some runtimes (notably Codex) do not reliably inherit shell exports; there, put bare KEY=value lines — no export, which docker --env-file rejects — in your own credential env file and load it with set -a; source /path/to/.env; set +a. Unprompted, the agent sources only the fixed user-owned locations (~/.tao/secrets.env, ~/.config/tao/.env); a repo-relative ./.env or /.env is loaded only when you point the agent at it. On session start the hook reports which vars it detects in the environment (names only). The agent never reads credential values — it only checks presence, and never writes a credential to a file you did not ask for; registry logins (docker login nvcr.io), the one-time enroot credentials install on a SLURM cluster (~/.config/enroot/.credentials, which persists your NGC key in plaintext on shared storage), and SLURM job sidecars happen only after you approve them.
When a workflow needs Hugging Face access, get a token from Hugging Face settings and accept the model or dataset license before launch.
If a readiness check reports a missing CLI, container image, backbone, or credential, the TAO skills can often install or stage the missing piece after you approve the action. Ask the agent to continue the original workflow after a blocker is resolved; it should rerun preflight and proceed from the same task.
Persisting secrets is your own responsibility. If you’d rather not re-export each session, keep them in an env file and either source it yourself or point the agent at it; a shell rc or a secrets manager works just as well. The agent never creates that file for you.
Do I need to install anything else?
Usually there is no up-front setup beyond the chosen platform prerequisites.
Model and data skills run with just docker run; platform skills add job
tracking, S3 I/O, and multi-node over their native CLI with no
nvidia-tao-sdk — see Execution: no SDK required.
Application workflows such as IAA DEFT may provision documented host-Python
dependencies into an isolated workspace environment, but only after the
workflow’s approval gate. AutoML search (tao-run-automl) remains the one
special SDK exception: its Preflight lazily installs the
nvidia-tao-automl wheel, whose pin lives in versions.yaml
(wheels.tao_automl_*).
Updating
Because the install pins a release tag, update / upgrade re-fetch that same
tag — they do not move you to a newer release. To move to a new release, check
the Releases page and
re-add the marketplace at the new tag; the existing entry is replaced in place.
Claude Code: re-point the marketplace, then update the plugin — /plugin install is a no-op on an already-installed plugin, so it will leave you on the
old build even after the marketplace moves.
/plugin marketplace add NVIDIA-TAO/tao-skill-bank@
Then update the installed plugin, either from /plugin manage in-session or
from a shell:
claude plugin update tao-skills@tao-skill-bank
Restart, or run /reload-plugins, to apply. To re-fetch the currently pinned
tag (e.g. after a cache wipe):
/plugin marketplace update tao-skill-bank
/reload-plugins
If skills look stale (cached contents):
rm -rf ~/.claude/plugins/cache/tao-skill-bank
then re-run /plugin install.
Codex:
# Move to a new release. Codex refuses to re-add a marketplace that is already
# registered at a different ref, so drop the old registration first.
codex plugin marketplace remove tao-local-plugins
codex plugin marketplace add NVIDIA-TAO/tao-skill-bank@
codex plugin add tao-skill-bank@tao-local-plugins
# Or, to re-fetch the tag you are already pinned to:
codex plugin marketplace upgrade tao-local-plugins
Note the marketplace is named tao-local-plugins (the name field in
.agents/plugins/marketplace.json); tao-skill-bank is the plugin name.
Re-running scripts/install-codex-agents.sh does all three steps for you.
If you copied AGENTS.md to ~/.codex/AGENTS.md, re-copy from the upgraded plugin cache to pick up identity changes.
Getting started (5 minutes)
The quickest way to verify your setup: run a Visual ChangeNet inference on a sample image.
Prerequisites
set -a; source /path/to/.env; set +a # omit if already exported
docker --version
docker run --rm --gpus all nvidia/cuda:12.2.0-base-ubuntu22.04 nvidia-smi
echo "$NGC_KEY" | docker login nvcr.io -u '$oauthtoken' --password-stdin
If any check fails, see skills/platform/tao-run-on-docker/SKILL.md for install/troubleshooting.
Smoke test
In a Claude Code session with the plugin installed, ask:
“Run Visual ChangeNet inference on this sample image: /tmp/sample.png. Write results to /tmp/vcn-out/.”
The agent will read skills/models/tao-train-visual-changenet/SKILL.md (skill name tao-train-visual-changenet, plus its references/skill_info.yaml if present), construct a docker run --gpus all ... invocation, and execute via Bash. No Python needed. No SDK install. Just docker + the plugin. For classify mode, expect per-image PASS/NO_PASS-style predictions and result files under /tmp/vcn-out/. For segment mode, expect binary change-mask outputs under the requested results directory.
For more complex workflows, see skills/applications/tao-run-deft-aoi/SKILL.md
(tao-run-deft-aoi, shorthand tao-deft-aoi) for AOI iterative fine-tuning,
skills/applications/tao-run-deft-iaa/SKILL.md (tao-run-deft-iaa, shorthand
tao-deft-iaa) for the self-contained local-Docker IAA loop, and
skills/applications/tao-run-automl/SKILL.md (tao-run-automl) for
hyperparameter optimization. AutoML launch reviews should show the number of
recommendations, metric, search space, expected runtime, and resolved train
image before long-running jobs start.
What’s in the bank
| Layer | Purpose | Examples |
|---|---|---|
skills/models/ |
Network-centric skills: containers, commands, data formats, checkpoints | tao-finetune-cosmos-reason, tao-train-visual-changenet, tao-finetune-clip, tao-train-dino, tao-train-segformer, … |
skills/data/ |
Data preparation, analysis, and enhancement | tao-mine-aoi-images, tao-analyze-gaps-visual-changenet, tao-route-visual-changenet-samples, tao-analyze-gaps-vlm-bcq, tao-convert-dataset-format, tao-validate-dataset-format, tao-generate-image-grounding, tao-generate-referring-expressions, tao-generate-video-reasoning-annotations |
skills/platform/ |
Where and how jobs run | tao-run-on-docker (local daemon or DOCKER_HOST=ssh://), tao-run-on-brev (instance-based GPU), tao-run-on-slurm (remote SLURM cluster), tao-run-on-kubernetes (k8s), tao-run-on-virtualenv (docker-free local venv), tao-data-io (S3/data staging), tao-setup-nvidia-gpu-host (host runtime) |
skills/applications/ |
End-to-end workflows composing the layers above | tao-run-deft-aoi, tao-run-deft-iaa, tao-run-automl-deft-pipeline, tao-analyze-changenet-rca, tao-train-single-step, tao-run-automl, tao-finetune-huggingface-model, tao-port-huggingface-model, tao-run-inference-service |
Each skill is a directory with SKILL.md (agent-readable instructions). Optional references/skill_info.yaml provides structured metadata (container image, per-action command/mode/inputs/outputs) the agent uses to construct the container command; optional scripts/ bundles supporting code.
The skills/core/ directory is not a second copy of the skill bank. It is the Codex plugin surface for small helper/router skills, such as capability discovery and launch intake. Canonical model, data, platform, and application skills live once in the layer directories above; do not add symlinks or copies under skills/core/.
Execution: no SDK required
Job tracking, S3 I/O, multi-node training, and failure-classified retries are
all built into the bank — there is no nvidia-tao-sdk. Every platform skill
implements the same four-verb consumer contract (submit/status/logs/
cancel) over its native CLI (docker, kubectl, ssh+sbatch, brev, or
the vendored virtualenv runner):
- Job tracking —
scripts/tao_job_record.pymints a job id and bindsresults_dirbefore launch (record-then-launch), then records state transitions in a fixed vocabulary. - S3 / data staging — the
tao-data-ioskill stages inputs (storage tiers A/B/C) and uploads results. - Multi-node — the SLURM/K8s multi-node templates in
templates/plus the NCCL probe (scripts/nccl_allreduce_probe.py). - Retries — infra-vs-program failure classification lives in
tao-launch-workflow(prose the agent applies, not a brittle regex table).
These need only the platform CLI plus the bank’s helper scripts — no wheel to
install. The one exception is AutoML hyperparameter search
(tao-run-automl), which uses the nvidia-tao-automl wheel (and its transitive
nvidia-tao-sdk dependency) to pick each next config; its Preflight installs
the right extra on first use. The pins live in versions.yaml
(wheels.tao_automl_*).
Contributing a new skill
Read first:
docs/skill-requirements.md— the must-follow rules for naming, frontmatter, size,evals/evals.json, and the security-scanner gotchas that block at signing. CI errors and signing-pipeline blockers are called out separately. Treat this as the authoritative gate list;docs/authoring.mdis the longer walkthrough.
See docs/authoring.md for the full authoring guide. The minimum viable skill is just SKILL.md — references/skill_info.yaml and friends are optional and only added when they earn their keep.
In brief:
- Pick the layer (
skills/models/,skills/data/,skills/platform/,skills/applications/). - Copy a template from
templates/skill-skeleton/—minimal/for the bare path,model/,data/,platform/, orworkflow/for richer scaffolding. - Fill in frontmatter and SKILL.md body. Body must contain a
## Quick Startsection, adocker runblock, or a link toreferences/skill_info.yaml. - Add
evals/evals.json(required for Tier-3 signing — seedocs/skill-requirements.md§ 2.3).eval.configis optional and only needed if you want live-execution coverage. - Add the skill path to
.claude-plugin/marketplace.jsonunder the relevant plugin(s). - Do not add a mirror entry under
skills/core/; Codex helper skills route to the canonical layer directories. - Validate with
scripts/validate-skills.shbefore submitting a PR.
Repository structure
tao-skills-external/
├── .claude-plugin/
│ ├── marketplace.json # marketplace catalog (plugin definitions)
│ └── plugin.json # plugin manifest (fallback when loaded directly)
├── hooks/
│ ├── hooks.json # SessionStart hook registration
│ └── session_start.sh # emits agent guidance; reports credential vars present in the env
├── .codex-plugin/
│ └── plugin.json # Codex plugin manifest
├── .agents/
│ └── plugins/marketplace.json # Codex marketplace entry
├── versions.yaml # single source of truth: container images + AutoML wheel versions
├── README.md
├── docs/
│ ├── skill-requirements.md # must-follow rules: naming + signing gates (read first)
│ ├── authoring.md # guide for adding new skills
│ └── maintenance.md # RC bump procedure for versions.yaml
├── templates/skill-skeleton/ # copy-paste starting points (minimal + per-layer)
├── scripts/
│ ├── validate-skills.sh # CI validator
│ ├── verify-standalone.sh # end-to-end smoke (docker-only path)
│ ├── install-codex-agents.sh # one-shot Codex install: marketplace + plugin + AGENTS.md
│ └── migrate-to-version-keys.py # one-shot: literal nvcr.io paths → versions.yaml keys
└── skills/
├── applications/ # 13 end-to-end workflow skills
├── data/ # 10 data preparation/analysis skills
├── models/ # 53 network-centric skills
├── platform/ # 7 compute backend / runtime skills
└── core/ # 2 Codex helper/router skills; no mirrored skill symlinks
CI
The repo runs three CI suites in parallel:
- NV-ACES skill evaluation (
.skill-eval.yml) — Tier 1/2 quality scoring, security scan. - Skill execution eval (
.gitlab-ci.yml) — runs each skill’seval.configon a real GPU runner. validate-skills(scripts/validate-skills.sh) — marketplace path resolution, noskills/core/mirrors, frontmatter, body has runnable info, no SDK leaks, hook references resolve.
PRs must pass all three before merge.
Design rules
- Docker-native first. Every model/data skill should be runnable with just
docker run+ the contents ofSKILL.md. Platform skills add tracking/staging/multi-node via the four-verb contract — no SDK. - Generic docker conventions live once in
skills/platform/tao-run-on-docker. Other skills defer to it for--gpus, NGC auth, mount patterns, data-root relocation, etc. - No SDK in the bank.
tao_sdk-specific imports andsdk.create_job/build_entrypointcalls are allowed only underskills/applications/tao-run-automl(it keeps thenvidia-tao-automlwheel + its transitive SDK).scripts/validate-skills.shenforces this. - Minimum-viable skill is
SKILL.mdonly. Addreferences/skill_info.yamlonly when multi-action structured metadata earns its keep. - One canonical location per skill. Model, data, platform, and application skills live only in their layer directories;
skills/core/is for Codex helper/router skills, not mirrored copies. - Prefer portability over cleverness. A skill that works across three coding agents is more valuable than a skill that works perfectly in one.
Recommended Tools
Try a different keyword or remove a filter.
Install
npx skillfish add nvidia-tao/tao-skill-bank