Agent skill

nemoclaw-user-configure-inference

Lists all inference providers offered during NemoClaw onboarding. Use when explaining which providers are available, what the onboard wizard presents, or how inference routing works. Changes the active inference model without restarting the sandbox. Use when switching inference providers, changing the model runtime, or reconfiguring inference routing. Connects NemoClaw to a local inference server. Use when setting up Ollama, vLLM, TensorRT-LLM, NIM, or any OpenAI-compatible local model server with NemoClaw.

Stars 18,977
Forks 2,322

Install this agent skill to your Project

npx add-skill https://github.com/NVIDIA/NemoClaw/tree/main/.agents/skills/nemoclaw-user-configure-inference

SKILL.md

NemoClaw User Configure Inference

Lists all inference providers offered during NemoClaw onboarding. Use when explaining which providers are available, what the onboard wizard presents, or how inference routing works.

Context

NemoClaw supports multiple inference providers. During onboarding, the nemoclaw onboard wizard presents a numbered list of providers to choose from. Your selection determines where the agent's inference traffic is routed.

How Inference Routing Works

The agent inside the sandbox talks to inference.local. It never connects to a provider directly. OpenShell intercepts inference traffic on the host and forwards it to the provider you selected.

Provider credentials stay on the host. The sandbox does not receive your API key.

Provider Status

Provider Status Endpoint type Notes
NVIDIA Endpoints Tested OpenAI-compatible Hosted models on integrate.api.nvidia.com
OpenAI Tested Native OpenAI-compatible Uses OpenAI model IDs
Other OpenAI-compatible endpoint Tested Custom OpenAI-compatible For compatible proxies and gateways
Anthropic Tested Native Anthropic Uses anthropic-messages
Other Anthropic-compatible endpoint Tested Custom Anthropic-compatible For Claude proxies and compatible gateways
Google Gemini Tested OpenAI-compatible Uses Google's OpenAI-compatible endpoint
Local Ollama Caveated Local Ollama API Available when Ollama is installed or running on the host
Local NVIDIA NIM Experimental Local OpenAI-compatible Requires NEMOCLAW_EXPERIMENTAL=1 and a NIM-capable GPU
Local vLLM Experimental Local OpenAI-compatible Requires NEMOCLAW_EXPERIMENTAL=1 and a server already running on localhost:8000

Provider Options

The onboard wizard presents the following provider options by default. The first six are always available. Ollama appears when it is installed or running on the host.

Option Description Curated models
NVIDIA Endpoints Routes to models hosted on build.nvidia.com. You can also enter any model ID from the catalog. Set NVIDIA_API_KEY. Nemotron 3 Super 120B, Kimi K2.5, GLM-5, MiniMax M2.5, GPT-OSS 120B
OpenAI Routes to the OpenAI API. Set OPENAI_API_KEY. gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.4-pro-2026-03-05
Other OpenAI-compatible endpoint Routes to any server that implements /v1/chat/completions. If the endpoint also supports /responses with OpenClaw-style tool calling, NemoClaw can use that path; otherwise it falls back to /chat/completions. The wizard prompts for a base URL and model name. Works with OpenRouter, LocalAI, llama.cpp, or any compatible proxy. Set COMPATIBLE_API_KEY. You provide the model name.
Anthropic Routes to the Anthropic Messages API. Set ANTHROPIC_API_KEY. claude-sonnet-4-6, claude-haiku-4-5, claude-opus-4-6
Other Anthropic-compatible endpoint Routes to any server that implements the Anthropic Messages API (/v1/messages). The wizard prompts for a base URL and model name. Set COMPATIBLE_ANTHROPIC_API_KEY. You provide the model name.
Google Gemini Routes to Google's OpenAI-compatible endpoint. NemoClaw prefers /responses only when the endpoint proves it can handle tool calling in a way OpenClaw uses; otherwise it falls back to /chat/completions. Set GEMINI_API_KEY. gemini-3.1-pro-preview, gemini-3.1-flash-lite-preview, gemini-3-flash-preview, gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite
Local Ollama Routes to a local Ollama instance on localhost:11434. NemoClaw detects installed models, offers starter models if none are present, pulls and warms the selected model, and validates it. Selected during onboarding. For more information, refer to Use a Local Inference Server (see the nemoclaw-user-configure-inference skill).

Experimental Options

The following local inference options require NEMOCLAW_EXPERIMENTAL=1 and, when prerequisites are met, appear in the onboarding selection list.

Option Condition Notes
Local NVIDIA NIM NIM-capable GPU detected Pulls and manages a NIM container.
Local vLLM vLLM running on localhost:8000 Auto-detects the loaded model.

For setup instructions, refer to Use a Local Inference Server (see the nemoclaw-user-configure-inference skill).

Validation

NemoClaw validates the selected provider and model before creating the sandbox. If validation fails, the wizard returns to provider selection.

Full details in references/inference-options.md.

Prerequisites

  • A running NemoClaw sandbox.
  • The OpenShell CLI on your PATH.
  • NemoClaw installed.
  • A local model server running, or Ollama installed. The NemoClaw onboard wizard can also start Ollama for you.

Change the active inference model while the sandbox is running. No restart is required.

Step 1: Switch to a Different Model

Switching happens through the OpenShell inference route. Use the provider and model that match the upstream you want to use.

NVIDIA Endpoints

console
$ openshell inference set --provider nvidia-prod --model nvidia/nemotron-3-super-120b-a12b

OpenAI

console
$ openshell inference set --provider openai-api --model gpt-5.4

Anthropic

console
$ openshell inference set --provider anthropic-prod --model claude-sonnet-4-6

Google Gemini

console
$ openshell inference set --provider gemini-api --model gemini-2.5-flash

Compatible Endpoints

If you onboarded a custom compatible endpoint, switch models with the provider created for that endpoint:

console
$ openshell inference set --provider compatible-endpoint --model <model-name>
console
$ openshell inference set --provider compatible-anthropic-endpoint --model <model-name>

If the provider itself needs to change, rerun nemoclaw onboard.

Step 2: Cross-Provider Switching

Switching to a different provider family (for example, from NVIDIA Endpoints to Anthropic) requires updating both the gateway route and the sandbox config.

Set the gateway route on the host:

console
$ openshell inference set --provider anthropic-prod --model claude-sonnet-4-6 --no-verify

Then set the override env vars and recreate the sandbox so they take effect at startup:

console
$ export NEMOCLAW_MODEL_OVERRIDE="anthropic/claude-sonnet-4-6"
$ export NEMOCLAW_INFERENCE_API_OVERRIDE="anthropic-messages"
$ nemoclaw onboard --resume --recreate-sandbox

The entrypoint patches openclaw.json at container startup with the override values. No image rebuild is needed. Remove the env vars and recreate the sandbox to revert to the original model.

NEMOCLAW_INFERENCE_API_OVERRIDE accepts openai-completions (for NVIDIA, OpenAI, Gemini, compatible endpoints) or anthropic-messages (for Anthropic and Anthropic-compatible endpoints). This variable is only needed when switching between provider families.

Step 3: Verify the Active Model

Run the status command to confirm the change:

console
$ nemoclaw <name> status

Add the --json flag for machine-readable output:

console
$ nemoclaw <name> status --json

The output includes the active provider, model, and endpoint.

Step 4: Notes

  • The host keeps provider credentials.
  • The sandbox continues to use inference.local.
  • Same-provider model switches take effect immediately via the gateway route alone.
  • Cross-provider switches also require NEMOCLAW_MODEL_OVERRIDE (and NEMOCLAW_INFERENCE_API_OVERRIDE) plus a sandbox recreate so the entrypoint patches the config at startup.
  • Overrides are applied at container startup. Changing or removing env vars requires a sandbox recreate to take effect.

NemoClaw can route inference to a model server running on your machine instead of a cloud API. This page covers Ollama, compatible-endpoint paths for other servers, and two experimental options for vLLM and NVIDIA NIM.

All approaches use the same inference.local routing model. The agent inside the sandbox never connects to your model server directly. OpenShell intercepts inference traffic and forwards it to the local endpoint you configure.

Step 5: Ollama

Ollama is the default local inference option. The onboard wizard detects Ollama automatically when it is installed or running on the host.

If Ollama is not running, NemoClaw starts it for you. On macOS, the wizard also offers to install Ollama through Homebrew if it is not present.

Run the onboard wizard.

console
$ nemoclaw onboard

Select Local Ollama from the provider list. NemoClaw lists installed models or offers starter models if none are installed. It pulls the selected model, loads it into memory, and validates it before continuing.

Linux with Docker

On Linux hosts that run NemoClaw with Docker, the sandbox reaches Ollama through http://host.openshell.internal:11434, not the host shell's localhost socket. If Ollama is already running, make sure it listens on 0.0.0.0:11434 instead of 127.0.0.1:11434.

console
$ OLLAMA_HOST=0.0.0.0:11434 ollama serve

If Ollama only binds loopback, NemoClaw can detect it on the host, but the sandbox-side validation step fails because containers cannot reach it.

Non-Interactive Setup

console
$ NEMOCLAW_PROVIDER=ollama \
  NEMOCLAW_MODEL=qwen2.5:14b \
  nemoclaw onboard --non-interactive

If NEMOCLAW_MODEL is not set, NemoClaw selects a default model based on available memory.

Variable Purpose
NEMOCLAW_PROVIDER Set to ollama.
NEMOCLAW_MODEL Ollama model tag to use. Optional.

Step 6: OpenAI-Compatible Server

This option works with any server that implements /v1/chat/completions, including vLLM, TensorRT-LLM, llama.cpp, LocalAI, and others. If the server also supports /v1/responses, NemoClaw only favors that path when onboarding can verify tool-calling behavior that matches what OpenClaw actually sends. Otherwise NemoClaw falls back to /v1/chat/completions.

Start your model server. The examples below use vLLM, but any OpenAI-compatible server works.

console
$ vllm serve meta-llama/Llama-3.1-8B-Instruct --port 8000

Run the onboard wizard.

console
$ nemoclaw onboard

When the wizard asks you to choose an inference provider, select Other OpenAI-compatible endpoint. Enter the base URL of your local server, for example http://localhost:8000/v1.

The wizard prompts for an API key. If your server does not require authentication, enter any non-empty string (for example, dummy).

NemoClaw validates the endpoint by sending a test inference request before continuing. For OpenAI-compatible endpoints, the validation prefers /responses only when the probe produces a compatible function or tool call. Endpoints that return 200 OK on /responses but do not format tool calls the way OpenClaw expects are configured to use /chat/completions instead.

Non-Interactive Setup

Set the following environment variables for scripted or CI/CD deployments.

console
$ NEMOCLAW_PROVIDER=custom \
  NEMOCLAW_ENDPOINT_URL=http://localhost:8000/v1 \
  NEMOCLAW_MODEL=meta-llama/Llama-3.1-8B-Instruct \
  COMPATIBLE_API_KEY=dummy \
  nemoclaw onboard --non-interactive
Variable Purpose
NEMOCLAW_PROVIDER Set to custom for an OpenAI-compatible endpoint.
NEMOCLAW_ENDPOINT_URL Base URL of the local server.
NEMOCLAW_MODEL Model ID as reported by the server.
COMPATIBLE_API_KEY API key for the endpoint. Use any non-empty value if authentication is not required.

Step 7: Anthropic-Compatible Server

If your local server implements the Anthropic Messages API (/v1/messages), choose Other Anthropic-compatible endpoint during onboarding instead.

console
$ nemoclaw onboard

For non-interactive setup, use NEMOCLAW_PROVIDER=anthropicCompatible and set COMPATIBLE_ANTHROPIC_API_KEY.

console
$ NEMOCLAW_PROVIDER=anthropicCompatible \
  NEMOCLAW_ENDPOINT_URL=http://localhost:8080 \
  NEMOCLAW_MODEL=my-model \
  COMPATIBLE_ANTHROPIC_API_KEY=dummy \
  nemoclaw onboard --non-interactive

Step 8: vLLM Auto-Detection (Experimental)

When vLLM is already running on localhost:8000, NemoClaw can detect it automatically and query the /v1/models endpoint to determine the loaded model.

Set the experimental flag and run onboard.

console
$ NEMOCLAW_EXPERIMENTAL=1 nemoclaw onboard

Select Local vLLM [experimental] from the provider list. NemoClaw detects the running model and validates the endpoint.

Note: NemoClaw forces the chat/completions API path for vLLM. The vLLM /v1/responses endpoint does not run the --tool-call-parser, so tool calls arrive as raw text.

Non-Interactive Setup

console
$ NEMOCLAW_EXPERIMENTAL=1 \
  NEMOCLAW_PROVIDER=vllm \
  nemoclaw onboard --non-interactive

NemoClaw auto-detects the model from the running vLLM instance. To override the model, set NEMOCLAW_MODEL.

Step 9: NVIDIA NIM (Experimental)

NemoClaw can pull, start, and manage a NIM container on hosts with a NIM-capable NVIDIA GPU.

Set the experimental flag and run onboard.

console
$ NEMOCLAW_EXPERIMENTAL=1 nemoclaw onboard

Select Local NVIDIA NIM [experimental] from the provider list. NemoClaw filters available models by GPU VRAM, pulls the NIM container image, starts it, and waits for it to become healthy before continuing.

Note: NIM uses vLLM internally. The same chat/completions API path restriction applies.

Non-Interactive Setup

console
$ NEMOCLAW_EXPERIMENTAL=1 \
  NEMOCLAW_PROVIDER=nim \
  nemoclaw onboard --non-interactive

To select a specific model, set NEMOCLAW_MODEL.

Step 10: Timeout Configuration

Local inference requests use a default timeout of 180 seconds. Large prompts on hardware such as DGX Spark can exceed shorter timeouts, so NemoClaw sets a higher default for local providers (Ollama, vLLM, NIM).

To override the timeout, set the NEMOCLAW_LOCAL_INFERENCE_TIMEOUT environment variable before onboarding:

console
$ export NEMOCLAW_LOCAL_INFERENCE_TIMEOUT=300
$ nemoclaw onboard

The value is in seconds. This setting is baked into the sandbox at build time. Changing it after onboarding requires re-running nemoclaw onboard.

Step 11: Verify the Configuration

After onboarding completes, confirm the active provider and model.

console
$ nemoclaw <name> status

The output shows the provider label (for example, "Local vLLM" or "Other OpenAI-compatible endpoint") and the active model.

Step 12: Switch Models at Runtime

You can change the model without re-running onboard. Refer to Switch Inference Models (see the nemoclaw-user-configure-inference skill) for the full procedure.

For compatible endpoints, the command is:

console
$ openshell inference set --provider compatible-endpoint --model <model-name>

If the provider itself needs to change (for example, switching from vLLM to a cloud API), rerun nemoclaw onboard.

Related Skills

  • nemoclaw-user-get-started — Quickstart for first-time installation

Expand your agent's capabilities with these related and highly-rated skills.

NVIDIA/NemoClaw

nemoclaw-user-skills-coding

Describes the agent skills shipped with NemoClaw and how to access them by cloning the repository. Use when users ask about AI agent support, coding assistant integration, or the .agents/skills/ directory.

18,977 2,322
Explore
NVIDIA/NemoClaw

nemoclaw-maintainer-day

Runs the daytime maintainer loop for NemoClaw, prioritizing items labeled with the current version target. Picks the highest-value item, executes the right workflow (merge gate, salvage, security sweep, test gaps, hotspot cooling, or sequencing), and reports progress. Use during the workday to land PRs and close issues. Designed for /loop (e.g. /loop 10m /nemoclaw-maintainer-day). Trigger keywords - maintainer day, work on PRs, land PRs, make progress, what's next, keep going, maintainer loop.

18,977 2,322
Explore
NVIDIA/NemoClaw

nemoclaw-user-manage-policy

Reviews and approves blocked agent network requests in the TUI. Use when approving or denying sandbox egress requests, managing blocked network calls, or using the approval TUI. Adds, removes, or modifies allowed endpoints in the sandbox policy. Use when customizing network policy, changing egress rules, or configuring sandbox endpoint access.

18,977 2,322
Explore
NVIDIA/NemoClaw

nemoclaw-user-overview

Explains how OpenClaw, OpenShell, and NemoClaw form the ecosystem, NemoClaw's position in the stack, what NemoClaw adds beyond the community sandbox, and when to prefer NemoClaw versus integrating OpenShell and OpenClaw directly. Use when users ask about the relationship between OpenClaw, OpenShell, and NemoClaw, or when to use NemoClaw versus OpenShell. Describes how NemoClaw works internally: CLI, plugin, blueprint runner, OpenShell orchestration, inference routing, and protection layers. Use for sandbox lifecycle and architecture mechanics; not for product definition (Overview) or multi-project placement (Ecosystem). Explains what NemoClaw covers: onboarding, lifecycle management, and management of OpenClaw within OpenShell containers, plus capabilities and why it exists. Use when users ask what NemoClaw is or what the project provides. For ecosystem placement or OpenShell-only paths, use the Ecosystem page; for internal mechanics, use How It Works. Lists changelogs and feature history for NemoClaw rel...

18,977 2,322
Explore
NVIDIA/NemoClaw

nemoclaw-maintainer-find-review-pr

Finds open GitHub PRs with security and priority-high labels, links each to its issue, detects duplicates (multiple PRs fixing the same issue), and presents a table of review candidates. Use when looking for the next PR to review. Trigger keywords - find pr, find review, next pr, pr to review, duplicate pr, security pr.

18,977 2,322
Explore
NVIDIA/NemoClaw

nemoclaw-maintainer-evening

Runs the end-of-day maintainer handoff for NemoClaw. Checks version target progress, bumps stragglers to the next patch version, generates a QA handoff summary, and cuts the release tag. Use at the end of the workday. Trigger keywords - evening, end of day, EOD, wrap up, ship it, cut tag, handoff, done for the day.

18,977 2,322
Explore

Didn't find tool you were looking for?

Be as detailed as possible for better results