agents365-ai/365-skills

bangumi-frames

Extract and organize frames from a Bilibili video (bangumi episode, UP upload, or a local file) into scenery shots and per-character image groups, using anime-specific person detection + CCIP character-identity embeddings.

Ver código-fonte
Documento original do Skill

Renderizado do repositório de origem, preservando títulos, exemplos, código, tabelas, links e imagens.

bangumi-frames — Bilibili Anime Frame & Character Organizer

Overview

Give a Bilibili video (a bangumi ep link, a UP-upload BV link/id, or a local video file); it downloads → extracts scene-change keyframes → splits scenery vs character frames → organizes the character crops. One pass, two modes:

  • no `--ref` (cluster mode) — group every character crop by CCIP identity into

characters/char_NN/.

  • with `--ref DIR` (one-vs-rest mode) — given ONE character's reference folder,

pull every crop in the video that matches it into matched/, filenames prefixed with distance (closest first) so a tight threshold yields a pure set.

Models are anime-specific (deepghs anime person detection + CCIP character-identity embeddings) — they do not work on live-action footage.

When to use / when NOT to use

  • Use when the user wants to collect/extract/organize anime frames or screenshots

from a Bilibili video — by character, by scenery, or to pull out one specific person.

  • Don't use for live-action video (needs an insightface-class face stack instead),

or for generic video editing/trimming/transcoding.

Bundled resources

ResourceRead it when
references/pipeline.mdTuning a stage — download (--height/--prefer), extract (--scene/--interval/--dedup/--skip), --clean (OCR+LaMa subtitle/watermark removal), classify (--conf/--min-area); feature caching; the CPU/CoreML rule; --redo
references/modes.mdChoosing/tuning the two modes — mode 1 cluster (--eps/--min-samples) vs mode 2 one-vs-rest (--ref-eps, the distance-band histogram, the compressed-embedding threshold lore); full output layout
scripts/bangumi_frames.pyThe entry point (all stages + both modes)
scripts/remove_overlay.pyStandalone subtitle/watermark removal on a frame dir or single image

Prerequisites

  1. ffmpeg on PATH; yt-dlp on PATH for downloads (a local-file input skips download).
  2. Python 3.9+, pip install dghs-imgutils (first run pulls ~300 MB of models from

HuggingFace, then cached locally).

  1. A Bilibili cookie (Netscape cookies.txt). Resolution order:

--cookies > $BILIBILI_COOKIES > ~/bb_up/bb_cookies/www.bilibili.com_cookies.txt. 1080p+ / premium episodes need a cookie with membership; a preview-only download means the cookie lacks access to that episode. Local-file input needs no cookie.

  1. Run the CCIP step on CPU — do not set ONNX_MODE=CoreML (CCIP crashes; the script

pops it before clustering/matching). Person detection is fine on CoreML.

  1. (Only for --clean) pip install rapidocr-onnxruntime simple-lama-inpainting.
  2. (Only for --engine pyscenedetect) pip install scenedetect.

Usage

bash
SKILL=skills/bangumi-frames/scripts/bangumi_frames.py

# Mode 1 — cluster everyone into char_NN groups
python3 $SKILL https://www.bilibili.com/video/BV15qVm68E2h --out ~/frames
python3 $SKILL ep1231575 --out ~/frames           # ep / BV id also accepted
python3 $SKILL ~/local.mp4 --out ~/frames          # local file, skips download

# Mode 2 — pull out ONE character (ref folder = ~200 crops of that character)
python3 $SKILL BV15qVm68E2h --ref ~/refs/紫灵 --ref-eps 0.04 --out ~/frames

# Optional: strip burned-in subtitles + watermark before analysis
python3 $SKILL ep1231575 --clean --out ~/frames

Stages are idempotent (a stage is skipped when its output already exists; clustering / matching always re-runs since the CCIP features are cached). For every flag, the per-stage trade-offs, and the threshold lore, read the two reference files above.

Agent-native output: stdout is a single JSON envelope ({"ok", "data", "next", "meta"} on success, {"ok": false, "error"} on failure — JSON when piped, a human summary on a TTY; force with --format), stderr carries human progress logs, and exit codes are stable (0 ok · 1 runtime · 2 auth · 3 validation). Use --dry-run to preview the plan without downloading, --schema to print the output contract. Details in references/pipeline.md.

Output

<out>/<id>/                     # id = BV id / ep id / local filename
├── frames/  frames.json        # keyframes + timestamps
├── scenery/                    # frames with no detected character
├── crops/  features.npy        # character crops + cached CCIP features
├── detect.json                 # frame -> person boxes / crops
├── characters/                 # MODE 1: char_NN_crop/ + char_NN_full/ (paired), _unsorted/, _montage.png
├── matched/                    # MODE 2: 0.012_<crop>.jpg (distance-prefixed) + index.json
├── matched_montage.png         # MODE 2 sample montage
└── index.json                  # MODE 1: char group -> {crop, frame, time}

After a run, look at characters/_montage.png (mode 1) or matched_montage.png (mode 2) first to judge quality, then read index.json. See references/modes.md for what to adjust when grouping/matching is off.

Limits

  • Anime / 2.5D-render art only; live-action needs a different (face-recognition) stack.
  • CCIP may split one character's different forms (outfit / transform) into separate groups

— usually fine for "group by visual appearance"; for mode 2, put each form in the ref.

  • 1080p+ on Bilibili needs a membership cookie; download is for personal offline analysis

only and uploads nothing.

do mesmo repositório

Mais Skills

Todos os Skills
agents365-ai
Comunidade

drawio-skill

Create, edit, synchronize, inspect, test, and publish editable draw.io diagrams. Use when the user explicitly requests draw.io/diagrams.net, needs a polished architecture, ERD, UML, sequence, C4, SysML, BPMN, network, swimlane, ML, or infrastructure diagram, wants code/IaC/SQL/OpenAPI/AsyncAPI/Protobuf converted into a diagram, or wants an existing diagram queried, reviewed, diffed, restyled, kept in sync, or made interactive. Prefer Mermaid/PlantUML elsewhere when the requested artifact is diagrams-as-code rather than an editable draw.io file.

instalações
9
GitHub Stars
48
Atualizado
11 de set.
agents365-ai
Comunidade

agent-native-design

Use when designing, reviewing, or refactoring a CLI that must serve AI agents alongside humans, or when converting an API or SDK into an agent-usable CLI interface.

instalações
3
GitHub Stars
48
Atualizado
11 de set.
agents365-ai
Comunidade

assetseeker

Search free commercial-use creative assets across multiple online sources — photos, illustrations, icons, video footage, music, sound effects, and fonts. Use when the user needs to find assets for videos, PPTs, articles, or any content creation. Trigger on phrases like "find me a photo of", "search for icons", "I need background music", "find video footage", "search fonts", "找一张图片", "搜素材", "帮我找图标/视频/音乐/字体". Covers Pexels, Unsplash, Pixabay, Iconify, Freesound, Google Fonts and more — with API-backed search where available.

instalações
3
GitHub Stars
48
Atualizado
11 de set.
agents365-ai
Comunidade

asta-skill

Domain expertise for Ai2 Asta MCP tools (Semantic Scholar corpus). Intent-to-tool routing, safe defaults, workflow patterns, and pitfall warnings for academic paper search, citation traversal, and author discovery.

instalações
3
GitHub Stars
48
Atualizado
11 de set.