huaweicloud/huaweicloud-skills

huawei-cloud-openviking-embedding-switch

Switch OpenViking's embedding model to a local llama-server (or any OpenAI-compatible embedding endpoint) running inside a bwrap sandbox managed by job-env-manager.

查看源码
仓库原始内容

按源仓库内容呈现,保留标题、案例、代码、表格、链接以及原文引用的演示图片。

OpenViking Embedding Model Switch

概述

Switch the embedding model used by OpenViking to a local llama-server or any OpenAI-compatible endpoint, with proper vectordb index rebuild and sandbox-safe restart.

⚠️ Single-purpose skill — all operations go through the job-env-manager REST API (http://127.0.0.1:8090). Never run openviking-server directly on the host.

OpenViking is an AI context database that uses vector embeddings for semantic search. Its embedding model is configured in ov.conf under the embedding.dense section. When switching to a different embedding model (especially one with a different vector dimension), the existing vectordb index must be deleted and rebuilt — otherwise OpenViking raises EmbeddingRebuildRequiredError on startup.

Architecture

OpenViking Embedding Model Switch
├── Detect current config     (Read ov.conf embedding.dense section)
├── Validate endpoint         (Check llama-server /v1/embeddings)
├── Modify ov.conf            (Update provider, model, api_base, dimension)
├── Delete vectordb index     (If dimension changed: rm -rf vectordb/context)
├── Restart server            (Kill + exec, NOT stop/start)
└── Verify                    (Health + PID + dimension + log check)
┌─────────────────────────────────────────────────────┐
│                    Host                              │
│                                                      │
│  ┌─────────────┐    REST API   ┌──────────────────┐ │
│  │  Agent       │─────────────▶│  job-env-manager  │ │
│  │  (this skill)│              │  :8090            │ │
│  └─────────────┘              └────────┬─────────┘ │
│                                        │            │
│         ┌──────────────────────────────┼──────┐    │
│         │  bwrap sandbox (openviking)   │      │    │
│         │                               ▼      │    │
│         │  ┌────────────────────────────────┐  │    │
│         │  │  openviking-server :1933       │  │    │
│         │  │  ├── ov.conf (embedding config)│  │    │
│         │  │  ├── vectordb/context/         │  │    │
│         │  │  └── viking/ (metadata)        │  │    │
│         │  └────────────────────────────────┘  │    │
│         └──────────────────────────────────────┘    │
│                                                      │
│         ┌──────────────────────────────────────┐    │
│         │  bwrap sandbox (llama)                │    │
│         │  ┌────────────────────────────────┐  │    │
│         │  │  llama-server :18200           │  │    │
│         │  │  --embeddings --model bge-...  │  │    │
│         │  └────────────────────────────────┘  │    │
│         └──────────────────────────────────────┘    │
│                                                      │
│  Both sandboxes use --share-net, so 127.0.0.1        │
│  endpoints are mutually reachable.                   │
└─────────────────────────────────────────────────────┘

Prerequisites

Prerequisite check: job-env-manager running ``bash curl -s http://127.0.0.1:8090/api/v1/envs/openviking | python3 -c "import sys,json; print(json.load(sys.stdin)['state'])"
  • job-env-manager running on http://127.0.0.1:8090
  • OpenViking environment deployed and running (state = running)
  • llama-server running at 127.0.0.1:{port} with --embeddings flag
  • curl and python3 available on the host
  • No AK/SK or Huawei Cloud credentials required

IAM Permission Policies

This skill operates on local bwrap sandboxes via the job-env-manager REST API and does not access Huawei Cloud services — no Huawei Cloud IAM policies required. Equivalent access controls are listed in references/iam-policies.md.

核心命令 (Core Workflow)

Task 1: Detect Current Configuration

bash
SANDBOX_DIR=$(curl -s http://127.0.0.1:8090/api/v1/envs/openviking \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['cwd'])")

Read ov.conf under the sandbox directory to get the current embedding.dense section (provider, model, dimension).

Task 2: Validate Target Embedding Endpoint

bash
curl -s http://127.0.0.1:${LLAMA_PORT}/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"${MODEL_NAME}","input":"test"}' \
  | python3 -c "import sys,json; d=json.load(sys.stdin); print(len(d['data'][0]['embedding']))"

If unreachable, STOP. The script auto-corrects the dimension if the specified value doesn't match the actual endpoint output.

Task 3: Modify ov.conf

Backs up ov.conf to ov.conf.bak before modifying. Updates the embedding.dense section:

FieldDescription
providerEmbedding provider name
modelModel name (e.g., bge-small-zh-v1.5)
api_keyAPI key for the endpoint (empty for local)
api_baseEndpoint URL (e.g., http://127.0.0.1:18200/v1)
dimensionVector dimension (auto-corrected from endpoint)

Task 4: Delete Incompatible vectordb Index

⚠️ Critical: If dimensions differ, rm -rf vectordb/context is required. Otherwise EmbeddingRebuildRequiredError on startup.

If dimension is unchanged, skip this step.

Task 5: Restart openviking-server Inside the Sandbox

⚠️ Pitfall: POST /envs/openviking/stop + start re-runs start.sh, which overwrites ov.conf with TokenHub credentials. Do not use stop/start.

Instead:

  1. Kill old process from host: kill $PID, then poll for port 1933 release (up to 10s). If SIGTERM doesn't release the port, escalate to kill -9.
  2. Clean up stale lock files: .openviking.pid and vectordb LOCK files.
  3. Start new server via exec API with --max-time 15:
bash
curl -s --max-time 15 -X POST http://127.0.0.1:8090/api/v1/envs/openviking/exec \
  -H 'Content-Type: application/json' \
  -d '{"cmd":["bash","-c","nohup /root/runtime/openviking/venv/bin/openviking-server --config /workspace/process_dir/ov.conf > /workspace/process_dir/openviking-server.log 2>&1 & sleep 2 && echo started"]}'

Task 6: Verify

  1. Health check with retry loop (up to 30s): polls GET /health every second until healthy=true or timeout
  2. PID change check: verifies the new server PID differs from the old one (detects port conflict false positives)
  3. Collection dimension check: reads collection_meta.json and confirms Dimension matches target
  4. Log error check: precise grep for Traceback|ERROR.*Application startup failed|EmbeddingRebuildRequiredError|DataDirectoryLocked (avoids false positives from "Retrying" info messages)
  5. Rollback on failure: if health check fails or PID unchanged, restores ov.conf.bak and exits with error

Parameter Confirmation

ParameterRequiredDescriptionExample
MODEL_NAMEYesEmbedding model namebge-small-zh-v1.5
LLAMA_PORTYesllama-server port18200
TARGET_DIMENSIONYesVector dimension (auto-corrected if wrong)512
bash
# Usage
bash scripts/switch-embedding-model.sh <model_name> <llama_port> <dimension>

Common Embedding Model Dimensions

ModelDimensionTypical Use
bge-small-zh-v1.5512Lightweight Chinese embedding
bge-large-zh-v1.51024High-quality Chinese embedding
bge-small-en-v1.5384Lightweight English embedding
bge-base-en-v1.5768General-purpose English embedding
Qwen3-Embedding-0.6B1024Qwen3 embedding (TokenHub default)

Verification

See references/verification-method.md for step-by-step checks and end-to-end acceptance criteria.

Quick verification:

bash
# 1. Server healthy
curl -s http://127.0.0.1:1933/health \
  | python3 -c "import sys,json; assert json.load(sys.stdin)['healthy']; print('OK')"

# 2. Collection dimension matches target
python3 -c "import json; d=json.load(open('${SANDBOX_DIR}/data/vectordb/context/collection_meta.json')); assert d['Dimension']==${TARGET_DIMENSION}; print('OK')"

# 3. No errors in log (precise pattern)
grep -ci "Traceback\|Application startup failed\|EmbeddingRebuildRequiredError\|DataDirectoryLocked" \
  "${SANDBOX_DIR}/process_dir/openviking-server.log"
# Expected: 0

Guardrails

See references/guardrails.md for the full rules. Key principles:

  • Always run through job-env-manager — never execute openviking-server directly on the host
  • Never use stop/start restartstart.sh overwrites ov.conf with TokenHub credentials
  • Validate before modify — the target endpoint must respond before any config change
  • Rollback on failureov.conf.bak is restored if verification fails

References

DocumentDescription
config-reference.mdov.conf embedding section field reference
guardrails.mdSafety rules: sandbox execution, restart sequence, rollback
iam-policies.mdEquivalent access controls (no Huawei Cloud IAM needed)
verification-method.mdStep-by-step verification for each workflow
related-commands.mdCommon job-env-manager and curl commands
acceptance-criteria.mdAcceptance criteria for a successful switch
troubleshooting.mdTroubleshooting for common failure scenarios
dataflow-diagram.mdMermaid data flow diagram
demo/example-input.jsonExample input for the switch workflow
来自同一仓库

更多 Skills

全部 Skills
huaweicloud
社区

huawei-cloud-billing-scout

Huawei Cloud BSS billing (not AWS/Azure/other clouds; refuses pricing quotes, real-name review, and any non-billing scope): balance, spend, attribution, reconciliation, coupons, stored-value cards, enterprise/partner billing. One-page briefing via hcloud. Use only when the user explicitly mentions 华为云 / Huawei Cloud / BSS and 余额/账单/对账/资源包/代金券/储值卡/企业或伙伴账务; refuses pay, renew, refund, delete.

安装量
242
GitHub Stars
20
最近更新
8月31日
huaweicloud
社区

huawei-cloud-cci-instance-management

Huawei Cloud CCI Container Instance Lifecycle Management Overview Manage Huawei Cloud CCI (Cloud Container Instance) full lifecycle using hcloud CLI (KooCLI). CCI is a serverless container service — no cluster management needed, just create a Namespace, define a Network, then deploy workloads directly. Architecture : hcloud CLI → CCI OpenAPI → Namespace / Network / Deployment / StatefulSet / Pod / EIPPool / Service / Ingress Constraints and Rules Security Rules Two step confirmation : All destructive operations (delete Namespace/Network/Deployment/StatefulSet/Pod/EIPPool) require explicit user confirmation — preview command, resource details, and risk warning first; execute only after user confirms. Credential security : Never expose AK/SK values in conversation, commands, or output. Only use hcloud configure list to check credential status (presence only). Prefer profile mode or environ

安装量
242
GitHub Stars
20
最近更新
8月31日
huaweicloud
社区

huawei-cloud-computing-query

Queries Huawei Cloud computing resources (ECS/BMS/IMS/AS), Covers ECS instances, flavors, keypairs, quotas, server groups, block devices, NICs, VNC console, BMS bare metal servers/flavors/quotas, IMS images/OS versions/members/quotas, and AS scaling groups/configs/policies/activity logs/lifecycle hooks/warm pools/quotas. No write operations. Use this skill when the user needs to query ECS instance details, list flavors, check BMS availability, browse images, or view auto-scaling group/policy status. Triggers: 弹性云服务器, ECS, 裸金属, BMS, 镜像, IMS, 弹性伸缩, AS, 伸缩组, 伸缩策略, 规格查询, flavor, instance, image, scaling.

安装量
242
GitHub Stars
20
最近更新
8月31日
huaweicloud
社区

huawei-cloud-find-skills

Invoke this skill to search, discover, browse, find and install any Huawei Cloud (华为云) agent skill.Triggers include: "华为云","华为云有什么skill","华为云相关skill","华为云agent skill 市场","华为云skill类目","explore Huawei Cloud skills","show Huawei Cloud skill categories","does a Huawei Cloud skill exist for...","which Huawei Cloud skills exist","搜索华为云技能","有没有管理ECS/OBS/RDS的skill","帮我找 XX 华为云skill","介绍 XX Skill 内容","华为云 XX Skill 具体做什么","安装华为云Skill".

安装量
242
GitHub Stars
20
最近更新
8月31日