runpod/runpod-plugins-official

runpod-mcp

- Manage Runpod infrastructure — pods, serverless endpoints, jobs, templates, network volumes, container-registry auth, GPU/CPU catalog, and billing — via the Runpod MCP server's structured tool calls.

Zobacz źródło
Oryginalny dokument Skill

Treść z repozytorium z zachowaniem nagłówków, przykładów, kodu, tabel, linków i obrazów.

Runpod MCP

The Runpod MCP server exposes Runpod's control plane as structured tool calls, so an MCP-capable agent can manage infrastructure without shelling out. It is the same Runpod REST API that runpodctl uses — pick MCP when its tools are connected (typed params, structured errors, no shell quoting).

For a multi-step job, read the worked example before calling tools. Tool calls are easy to issue and easy to issue in the wrong order — the verified end-to-end sequences live in runpod/golden-paths/README.md (image → template → endpoint, pod → volume → serverless, multi-region, autoscaling, monitoring). This skill covers what each tool does; the paths cover what order to do them in and what it costs.

Connect

Connect the hosted server with your API key as a Bearer header if you also use runpodctl/flash — that one key auths the MCP and the CLIs (the 80% path):

bash
claude mcp add --transport http runpod -s user https://mcp.getrunpod.io/ \
  --header "Authorization: Bearer $RUNPOD_API_KEY"

Plain OAuth ("Sign in with Runpod", via npx @runpod/mcp-server@latest add) is MCP-only — the CLIs stay unauthed, so use it only for MCP-only work. Local stdio runs the server as a subprocess with your key. Those variants + the key-vs-OAuth tradeoff: [reference/connect.md](reference/connect.md). After connecting, reconnect the client (in Claude Code, /mcp) so the tools load.

Verify it's live (do this before relying on MCP): in Claude Code run /mcprunpod should show Connected, not Needs authentication (if it's the latter, sign in there first; the bundled plugin server registers the URL but stays inert until you authenticate). Confirm a real call works by asking for list-endpoints. If the runpod tools aren't present at all, the server isn't connected — (re)run the install above, or fall back to runpodctl for this task.

Check the server version (which REST API it drives): the MCP initialize handshake returns it in serverInfo.version. /mcp in Claude Code shows it, or probe the hosted server directly:

bash
printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"0"}}}' \
| curl -s -X POST https://mcp.getrunpod.io/ -H "Content-Type: application/json" \
    -H "Accept: application/json, text/event-stream" -H "Authorization: Bearer $RUNPOD_API_KEY" -d @-
# → serverInfo.version e.g. "3.0.0 [RUNPOD_REST_VERSION=v2]"  (verified 2026-07-29)

The MCP server drives Runpod's REST v2 internally (RUNPOD_REST_VERSION=v2), so most tools avoid the buggy public `rest.runpod.io/v1` control API. Two exceptions worth knowing: the Hub, public-endpoint and set-endpoint-gpus tools go through GraphQL (so they work under either REST version), and CPU serverless endpoints are not creatable through MCP — v2 has no CPU-endpoint concept at all (create-endpoint requires gpuPoolIds), so use runpodctl serverless create --compute-type CPU for those.

Prefer MCP or `runpodctl` over hand-rolled `rest.runpod.io/v1` calls for creating endpoints.

Tool surface

Structured tools, grouped by resource:

  • Pods — list, get, create, update, start, stop, restart, delete, stream logs.
  • Serverless endpoints — list, get, create, update, delete; list workers; list releases; stream worker logs.
  • Logs are no longer an MCP-only capability — runpodctl grew pod logs and serverless logs in v2.10.0. MCP still returns already-parsed, bounded frames, which is the easier shape inside an agent; reach for the CLI when you are shell-only or want --follow. Job output streaming (stream-job) remains MCP-only.
  • create-endpoint takes endpointType: QUEUE (default) or LOAD_BALANCER — see golden path 14. The routing type is fixed at creation; update-endpoint cannot change it.
  • Read an endpoint's invoke URLs from requestUrls on the get/list reply instead of assembling them.
  • To pin a specific GPU SKU on an existing endpoint use set-endpoint-gpus; create-endpoint/update-endpoint expose only gpuPoolIds and can't express a SKU (deploy-hub-repo can pin one at deploy time via gpuIds exclusions).
  • Jobs (serverless runtime) — run, runsync, status, stream, cancel, retry, health, purge queue.
  • Hublist-hub-repos (public catalog of prebuilt Serverless workers and Pod templates: vLLM, ComfyUI, …) and deploy-hub-repo, which deploys a repo's listed release as an endpoint — the same as clicking Deploy on the Hub.
  • Public endpointslist-public-endpoints: managed pay-per-use model APIs (text/image/video/audio) that need no deployment. Call the returned endpointId with run-endpoint/runsync-endpoint.
  • Templates — list, get, create, update, delete.
  • Network volumes — list, get, create, update, delete. create-network-volume takes volumeType (STANDARD | HIGH_PERFORMANCE) and a size of 10–4096 GB; omit volumeType to get the data center's default tier. The tier is immutable after creationupdate-network-volume can't change it.
  • Container registry auth — list, get, create, delete. A username + password for any registry; pass the resulting id as containerRegistryAuthId on create-pod/create-endpoint.
  • ECR delegations (list-/create-/delete-registry-delegation) — AWS ECR only, v2 only, and stores no credentials: you register a repository ARN and Runpod gets scoped pull access instead. Prefer it over a stored username/password for ECR. The reply carries a dockerRegistryUri — that's the image URI to deploy with.
  • Catalog — list/get GPU types, list/get CPU types, list/get data centers.
  • Billing — scoped usage/cost breakdowns (get-billing).
The tool list above is a map, not a contract. The server is the source of truth — /mcp (or your client's tool list) shows exactly what the connected version exposes, and each tool carries its own parameter descriptions. Check there before assuming a capability exists or doesn't.
Delete tools (delete-template, delete-pod, …) can return isError: true with "Unexpected end of JSON input" even on success — the Runpod REST API returns 204 No Content. Don't treat it as failure; confirm with a follow-up get-/list- (a deleted resource then 404s).

Use MCP vs runpodctl

  • Use runpod-mcp when the tools are connected AND the task is infra CRUD or a

serverless job call the server exposes. Cap large job/log output to a file.

  • Use runpodctl instead for: `send`/`receive` file transfer, SSH key

management, `doctor` setup, model cache — or any shell-only agent, or when the user wants a reproducible command.

  • Hand pod creation to runpodctl for a multi-GPU priority list (MCP's v2

create-pod takes one GPU type; extra gpuTypeIds are dropped with a _warning on success), or for a template + CPU pod together — create-pod rejects that combination, since a template deploy is GPU-and-v2-only. Each alone is fine in MCP: templateId (v2-only, imageName then optional, and each field you pass replaces the template's whole value rather than merging) or computeType: "CPU".

  • Not this lane: writing/deploying your own Python (→ flash); downloading

models or building/pushing images (→ companion-clis).

For concepts (pods vs serverless, GPU selection, storage), read ../runpod-usage/.

Source & docs

  • Server source: https://github.com/runpod/runpod-mcp
  • Package (npm): https://www.npmjs.com/package/@runpod/mcp-server
  • Hosted endpoint: https://mcp.getrunpod.io/
  • Docs: https://docs.runpod.io
z tego samego repozytorium

Więcej Skills

Wszystkie Skills
runpod
Społeczność

flash

- runpod-flash — code-first serverless: write Python locally, run it on remote Runpod GPUs/CPUs with flash dev (hot-reload + live worker logs), then flash deploy. Use for @Endpoint/@remote functions, resource config, and debugging flash deployments. For CLI-only infra management use runpodctl or runpod-mcp.

instalacje
8
GitHub Stars
46
Aktualizacja
21 wrz
runpod
Społeczność

runpod

- Start here for any Runpod task — running GPU/CPU pods, deploying serverless endpoints, templates, network volumes, building images, or understanding how Runpod works. Routes to the right skill (runpod-mcp, runpodctl, flash, companion-clis, runpod-usage, runpod-templates, runpod-migrate) and indexes two dozen live-verified end-to-end examples (golden paths) — use it when the lane is unclear, and for any multi-step or provisioning task even when it is not.

instalacje
8
GitHub Stars
46
Aktualizacja
21 wrz
runpod
Społeczność

runpod-migrate

- Migrate a codebase from the Runpod GraphQL API or REST v1 to REST v2 — inventory which parts use which API version, rewrite the call sites, flag breaking changes, and verify. Use when someone asks to move to v2, asks what v2 would change, or asks which Runpod API their code is on. For managing infrastructure rather than migrating code, use runpod-mcp or runpodctl.

instalacje
8
GitHub Stars
46
Aktualizacja
21 wrz
runpod
Społeczność

runpodctl

- Runpod CLI for managing GPU/CPU workloads from the terminal — pods, serverless endpoints, templates, network volumes, Hub deploys, models, SSH, and file transfer (send/receive). Use for terminal/CI/scripting, Hub browse/deploy, SSH setup, doctor, or when the Runpod MCP tools are not connected. For structured tool calls in an MCP-enabled session, prefer runpod-mcp.

instalacje
8
GitHub Stars
46
Aktualizacja
21 wrz