# AiVideoStudio **Repository Path**: osheyx/ai-video-studio ## Basic Information - **Project Name**: AiVideoStudio - **Description**: AI视频制作本地工具 - **Primary Language**: Unknown - **License**: MIT - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 1 - **Forks**: 0 - **Created**: 2026-07-01 - **Last Updated**: 2026-07-10 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # AIVideoStudio [![CI](https://github.com/kuaishou/ai-video-studio/actions/workflows/ci.yml/badge.svg)](https://github.com/kuaishou/ai-video-studio/actions/workflows/ci.yml) [![Docker](https://github.com/kuaishou/ai-video-studio/actions/workflows/release.yml/badge.svg)](https://github.com/kuaishou/ai-video-studio/actions/workflows/release.yml) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE) [![Python 3.11+](https://img.shields.io/badge/python-3.11%20%2F%203.12-blue.svg)](https://www.python.org/) [![Code style: ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff) **Local-first AI image / video production studio.** One closed loop that takes you from **generate -> review -> publish -> metrics -> next-prompt**, with built-in adapters for [ComfyUI](https://github.com/comfyanonymous/ComfyUI), the major cloud APIs, and a browser-based desktop shell. > See `docs/00_architecture.md` for the Chinese architecture doc, `docs/01_operations.md` > for ops, and `PRODUCTION_GUIDE.md` for production deployment. --- ## Table of Contents 1. [Highlights](#highlights) 2. [Quickstart](#quickstart) 3. [Architecture at a Glance](#architecture-at-a-glance) 4. [Project Layout](#project-layout) 5. [AI Editing Guide](#ai-editing-guide) 6. [The 8-Layer Rule (read this before editing)](#the-8-layer-rule-read-this-before-editing) 7. [Code Conventions for AI Assistants](#code-conventions-for-ai-assistants) 8. [Common Tasks](#common-tasks) 9. [Operations Cheat-Sheet](#operations-cheat-sheet) 10. [Documentation Map](#documentation-map) 11. [Contributing](#contributing) --- ## Highlights - **One studio, six modules**: understand -> viral-analysis -> produce -> feedback -> danmaku-commentary -> remix/highlights -- all wired into a single closed loop - **Provider-pluggable**: ComfyUI is the default local workflow provider; Seedream / Kling / Runway / Firefly slot in as HTTP adapters - **Local-first by default**: assets, state, and reports live on disk; switch the storage backend to any S3-compatible service when you scale - **Auto-closed-loop**: every published asset feeds back into the next prompt via Bayesian learning -- no manual copy-paste - **Automation-friendly**: drop JSON envelopes into `studio_storage/inbox/` from Hermes / OpenClaw; the watcher writes results to `outbox/` - **Desktop-friendly**: ships with a Python desktop shell (browser-as-UI); Tauri / Electron upgrade path is documented - **Container-aware editing runtime**: highlight / shot / remix / style / ffmpeg / watermark capabilities run as `CapabilityContainerSpec` with CPU + memory budgets applied at container startup (see `docs/21_editing_container_runtime.md`) --- ## Quickstart Requires **Python 3.11 or 3.12**. ```bash # 1. Clone & enter the repo git clone https://github.com/kuaishou/ai-video-studio.git cd ai-video-studio # 2. Set up the dev environment python -m venv .venv && source .venv/bin/activate pip install -e ".[dev]" # 3. Smoke-test the install pytest tests/ -q # unit suite, ~10s on a laptop aivs maintenance doctor # health check # 4. Run the API + desktop shell aivs api serve --host 127.0.0.1 --port 8080 # FastAPI aivs desktop serve --api http://127.0.0.1:8080 --port 8765 # static shell ``` Open `http://127.0.0.1:8765` for the dashboard (default landing: **量产中心**). ### Docker Quickstart ```bash docker compose -f docker/docker-compose.yml up -d # API on http://localhost:8080 # Web on http://localhost:5173 ``` For Kubernetes, see `deploy/helm/aivs/` and `PRODUCTION_GUIDE.md`. --- ## Architecture at a Glance ``` Entry points: CLI / FastAPI / Refine Web / Hermes / OpenClaw | MonetizationOrchestrator (business orchestration) | +----------+----------+ Capability Registry (15 atomic capabilities) | | | 6 modules: | | 1. Understand | | 2. Viral | <---- Bayesian | 3. Produce | learning | 4. Feedback | | 5. Danmaku | | 6. Remix | | | +----------+----------+ | Domain models + capability adapters | Infrastructure: Supabase / Redis / Kafka / RustFS | Cross-cutting: multi-tenant / metering / audit ``` Full diagram and layered walkthrough: `docs/09_architecture_overview.md`. --- ## Project Layout ``` ai_video_studio/ # Python package -- domain core + API + CLI |-- studio/ # Pure-Python domain core (no web framework) | |-- __init__.py # Layered architecture overview | |-- _layers.py # Frozen Layer / LayerRole constants | |-- reports.py # Shared helpers for /reports/* endpoints | |-- models.py # StudioTask / StudioAsset / Publication / ... | |-- orchestrator.py # submit / run / run_next / ingest / batch | |-- loop.py # Review / Publish / Analytics / PromptStrategy | |-- repository.py # JSON-backed state store | |-- inbox.py # Hermes / OpenClaw envelope schema + watcher | |-- config.py # StudioConfig + YAML + ${ENV} interpolation | |-- learning.py # Bayesian prompt-strategy learner | |-- event_bus.py # In-process pub/sub for cross-module hooks | `-- / # One folder per business module | |-- __init__.py # Public surface (Service class) | |-- domain.py # L4 dataclasses | |-- service.py # L3 orchestration | |-- llm_provider.py # L5 LLM adapter (pluggable) | `-- capabilities/ # L5 atomic capabilities (finest granularity) |-- api/ | |-- _deps.py # Shared FastAPI dependencies (orchestrator()) | |-- _helpers.py # Small helpers reused across routers (dedupe_ids, etc.) | |-- main.py # FastAPI app factory | `-- routers/ # Thin adapters -- one file per feature area | |-- studio.py # /api/v1/studio/* (task + report endpoints) | |-- studio_batch.py # /api/v1/studio/batch* + bulk ops | `-- ... # 25+ feature routers |-- cli/ # Typer-based aivs CLI (thin adapter) | |-- main.py # Aggregates sub-commands | |-- harness.py # aivs harness run -- closed-loop smoke runner | |-- maintenance.py # aivs maintenance {doctor,retry,validate-config} | `-- ... |-- workers/ # Async background workers (inbox / pipeline / metrics) `-- infra/ # Cross-cutting infra: cache, bus, containers, storage desktop/ # Browser-as-UI shell + reverse proxy web/ # Optional Refine / React dashboard deploy/helm/aivs/ # Kubernetes Helm chart docker/ # Dockerfiles + docker-compose.yml docs/ # Long-form architecture, ops, and feature docs tests/ |-- unit/ # Fast unit tests (no I/O) |-- integration/ # End-to-end HTTP tests (TestClient + tmp_path) `-- conftest.py # Shared fixtures (auto-disables SaaS gateway) ``` --- ## AI Editing Guide > **You are an AI assistant editing this repo.** Read this section first -- > it tells you where to look, what not to break, and the patterns to follow. **First, identify which layer you are in.** Look at the import paths and find the row that matches in [The 8-Layer Rule](#the-8-layer-rule-read-this-before-editing). Most edits to a `bulk_*` endpoint or a `/tasks/...` route live in L1 (`api/routers/`); new business logic in L3 (`studio//`); new shared helpers in L5/L6 depending on what they touch. **Before writing code, search the helpers.** Open [`studio/reports.py`](ai_video_studio/studio/reports.py) and [`api/_helpers.py`](ai_video_studio/api/_helpers.py) first. In particular, *any* new `bulk_*` route should reuse `bulk_apply_to_tasks` -- the table in [Shared helpers](#shared-helpers-do-use-dont-reimplement) lists every helper and what it replaces. **The closed loop is sacred.** This repo exists to wire `understand -> viral-analysis -> produce -> feedback -> danmaku-commentary -> remix/highlights` into a single loop. When you add a route, ask: *does this feed the next stage?* If not, it is probably a CLI command or a helper, not a new HTTP endpoint. **Test before you push.** Every public route has an integration test under [`tests/integration/test_v4_e2e_http.py`](tests/integration/test_v4_e2e_http.py) using `_running_servers` + `_http`. Pattern to copy: ```bash AIVS_SAAS_GATEWAY=0 .venv/bin/pytest tests/integration/test_v4_e2e_http.py -k "bulk_xxx" -q ``` **Don't trust comments about line counts.** This README was last regenerated at 469 lines; the source files evolve. When in doubt, run `wc -l ai_video_studio/api/routers/studio_batch.py` to see the current size. **Slow down on tests.** The full integration suite is 1400+ tests and takes ~10 minutes wall-time on a workstation. Run the smallest `-k` selection that covers your change, then expand. --- ## The 8-Layer Rule (read this before editing) This is the single most important rule in the codebase. Everything else (folder layout, module naming, where to put new code) follows from it. The rules are also **enforced in code** via `ai_video_studio.studio._layers` -- see `assert_layer_dependency` for runtime guards. | Layer | Path | Purpose | Allowed deps | | --- | --- | --- | --- | | **L1 Entry** | api/, cli/, desktop/, workers/ | HTTP / CLI / worker adapters | L2 only | | **L2 Orchestration** | studio/orchestrator.py, studio/monetization/ | Cross-module business flow | L3, L4 | | **L3 Business module** | studio/{understanding,analysis,production,feedback,commentary,remix}/ | One module = one capability area | L4, L5 | | **L4 Domain model** | studio/domain.py, studio/commentary_domain.py | Dataclasses + invariants | (no deps) | | **L5 Capability adapter** | studio/llm/, studio/media/, /llm_provider.py, /capabilities/ | Pluggable LLM / media backends | L6 | | **L6 Infrastructure** | infra/ | Storage / cache / bus / containers | (leaf) | | **L7 Cross-cutting** | studio/{safety,resilience,observability,enterprise}/ | Auth, retries, tracing, multi-tenant | Anywhere (use sparingly) | | **L8 Config & deploy** | studio/config.py, studio/deploy/ | Startup wiring | (read-only everywhere) | ### Hard rules 1. **L1 must never import L5 directly.** API routes orchestrate via L2/L3 services. 2. **L4 is leaf-only.** Dataclasses do not import from any other layer. 3. **L5 capabilities must be atomic and side-effect-aware.** Each capability takes a typed request, returns a typed result, and lives in a registry. 4. **Anything stateful goes through `repository.py`** (JSON files under `studio_storage/`). No module owns its own persistence. 5. **New business area = new folder under `studio/`.** Follow the existing layout (`/__init__.py`, `/domain.py`, `/service.py`, `/capabilities/`). ### Where do I put my code? | I am adding... | Put it in... | | --- | --- | | A new HTTP endpoint | `api/routers/.py` (delegate to a service) | | A new CLI sub-command | `cli/_cmd.py` + register in `cli/main.py` | | A new business module | `studio//` (full folder skeleton) | | A new atomic capability | `studio//capabilities/.py` + register in `registry.py` | | A new LLM provider | `studio//llm_provider.py` + register in `__init__.py` | | A new cross-cutting concern | `studio//` | | A new shared helper | The lowest layer that needs it (avoid `api/` for shared code) | --- ## Code Conventions for AI Assistants > If you are an AI editing this codebase: **read this section before touching > any file**. These rules exist because the codebase is large (65k+ lines, 400+ > files) and small inconsistencies compound quickly. ### Style - **Python 3.11+** only. `from __future__ import annotations` is kept for compatibility with existing modules; new files do not need it. - **Formatter & linter:** `ruff` (configured in `pyproject.toml`). Run `make lint` before committing. Auto-fix with `ruff check --fix`. - **Naming:** modules `snake_case`; classes `PascalCase`; functions `snake_case`; Pydantic models `PascalCase` ending in `Request`, `Response`, `Create`, or `Update`. - **Comments in code:** the codebase is bilingual (zh / en). Prefer zh for design rationale and en for technical keywords. Docstrings are English. - **No bare `except`.** Catch the specific exception you handle. - **No global mutable state outside `repository.py`.** All state goes through the file-backed repository so the next SQLite/MySQL backend can be a drop-in. ### Function length - **Target <= 30 lines** for any function. Anything over 60 lines almost always has an extractable helper. - **If you find yourself reaching for `time`, `re`, `collections.Counter` inside a router function, stop and look in `studio/reports.py` first.** That module hosts shared helpers used across `/reports/*` endpoints. ### Shared helpers (DO use, DON'T reimplement) | Helper | Purpose | Replaces | | --- | --- | --- | | `studio.reports.clamp_int(value, low=, high=)` | Bound query parameters | `max(lo, min(hi, int(x)))` | | `studio.reports.clamp_days(value)` | Clamp `days` to `[1, 90]` | same as above for `days` | | `studio.reports.recent_visible_tasks(tasks, days=, now=)` | Filter by time window + archived/pinned rule | 3-line list comprehension | | `studio.reports.prompt_templatize(prompt)` | Replace numbers / long tokens / Chinese runs with placeholders | regex soup | | `studio.reports.words(prompt)` | Lowercase English tokens for prompt analytics | `re.findall(...)` boilerplate | | `studio.reports.STOPWORDS` | The shared English stop-word set | A local 80-line `stop = {...}` block | | `studio.reports.NgramTally` | Accumulate pattern -> (count, tasks, batches, sample) | Triple-counter dict juggling | | `api._deps.orchestrator()` | Build a fresh `StudioOrchestrator` | `StudioOrchestrator(load_studio_config())` in every router | | `api._helpers.dedupe_ids(ids)` | Remove duplicate task ids while preserving order | 7-line `seen`/`unique_ids` set dance | | `studio.reports.summarize_batch(tasks, ...)` | Per-batch aggregation (status / duration / stuck / top prompts / tags) | 90-line inline loop | | `studio.reports.group_by_batch_id(tasks)` | Group tasks by `metadata.batch_id` | 4-line `setdefault` loop | | `studio.reports.bucket_key(ts, bucket=)` | Format a timestamp into "YYYY-MM-DD HH:00" or "YYYY-MM-DD" | inline `datetime.strftime` blocks | | `studio.reports.aggregate_by_time_bucket(tasks, bucket=)` | Group tasks into time buckets with status counts | 25-line inline defaultdict loop | | `studio.reports.bucket_rows(grouped)` | Convert grouped bucket dicts into JSON rows + totals | 20-line inline sort+aggregate loop | | `studio.reports.task_cost_factor(status)` | Map status -> cost factor (done=1.0, failed=0.5, else=0.0) | 3-branch if/elif/else | | `studio.reports.clone_source_row(repo, tasks_by_id, src, events)` | Per-source aggregation for /reports/clone-fields-effectiveness | 25-line inline loop | | `studio.reports.jaccard(a, b)` | Jaccard similarity between two sets | 4-line edge-case handling | | `studio.reports.cluster_by_similarity(items, threshold=)` | Greedy agglomerative clustering by set similarity | 15-line nested loop | | `studio.reports.task_status_value(task)` | Get ``task.status.value`` (handles both enum and string) | `s = t.status.value if hasattr(t.status, "value") else t.status` | | `studio.reports.count_by_status(tasks)` | Bucket tasks by status and return counts | `done = sum(1 for t in items if t.status.value == "done"); ...` per status | | `studio.reports.group_by_prompt(tasks, *, include_assets=, ...)` | Group tasks by stripped prompt with status counts + asset counts | 25-line `defaultdict(_new_agg)` + for-loop in top-prompts reports | | `studio.reports.BulkDecision(kind, log_key, log_entry, ...)` | Per-task decision for `bulk_apply_to_tasks` (`"update"` / `"skip"` / `"error"`) | inline `if decision == "skip": ...` branching | | `studio.reports.BulkApplyResult(updated, skipped_locked, skipped_other, errors, affected_batches)` | Raw outcome buckets from `bulk_apply_to_tasks` | four parallel `list` accumulators in every bulk route | | `studio.reports.bulk_apply_to_tasks(orch, ids, *, include_locked, mutate, now=None, dry_run=False)` | Iterate ids, call `mutate(task)`, handle existence / locked / save / log | the 40-line loop in every `bulk_*` route | ### Bulk endpoint pattern Every endpoint that does `for t in tasks: mutate(t); save(t); log; count` is a candidate for `bulk_apply_to_tasks`. The shape is: ```python async def bulk_xxx(req: XxxRequest) -> dict[str, Any]: orchestrator = _deps.orchestrator() now = time.time() def mutate(task): # Decide + apply the field change in place. # Return one of: return _reports.BulkDecision( kind="update", log_key="bulk_xxx_log", log_entry={"old": old, "new": new, "by": req.by, "at": now}, extras={"row": "for response"}, ) # OR kind="skip", skip_detail={"reason": "no_change"} # OR kind="error", error_message="..." result = _reports.bulk_apply_to_tasks( orchestrator, req.task_ids, include_locked=req.include_locked, mutate=mutate, now=now, ) # Map buckets onto the response shape (each endpoint exposes different skips). return {"updated": result.updated, "skipped_locked": result.skipped_locked, ...} ``` Rules of thumb: - **Outer filter** (status / tag / batch_id / created_after) is the *endpoint's* responsibility, not the helper's; the helper only handles existence + locked + save + log. - **Per-bucket skip reasons** (e.g. `no_change`, `already_has_tag`, `overwrite_protected`) live in `BulkDecision.skip_detail["reason"]` -- the endpoint reads `result.skipped_other` and filters to its own buckets. - **Always set `now=time.time()` once at the top** so all log entries share a consistent timestamp. - **Pre-compute closure variables** (`tag_set`, `new_prompt`, `by`, ...) before defining `mutate`; don't read request fields from inside the closure. - **`dry_run=True`** (helper parameter) runs `mutate` and populates `result.updated` but skips `save_task` and the `updated_at` bump -- the row's `extras` / `log_entry` still appear in the response so callers can preview. ### Endpoint rules 1. **Thin adapters.** A router function should: parse input -> call one service method -> return `dataclass.to_dict()`. If you find yourself writing business logic in a router, move it into a service. 2. **Lazy `load_studio_config()` imports inside functions** are intentional in this codebase (to avoid import-time side effects). Do not move them to the top of the file unless you know the file has no circular-import risk. 3. **Hidden state assumption:** "hide archived unless pinned" is the default for every `/reports/*` endpoint. Use `recent_visible_tasks` to enforce it. 4. **Pydantic models** for request bodies, plain dataclasses for responses (call `.to_dict()` to serialize). ### Test rules 1. **Unit tests** in `tests/unit/` -- no I/O, no network, fast. 2. **Integration tests** in `tests/integration/` -- use `TestClient` and `tmp_path` for isolation. `tests/conftest.py` auto-sets `AIVS_SAAS_GATEWAY=0` so the SaaS auth layer is bypassed. 3. **Run `pytest tests/unit/ -q` first** during development. Integration tests take longer; run them before opening a PR. 4. **Do not delete or weaken a test to make it pass.** If a test fails, either the change is wrong, or the test was asserting implementation detail -- surface it in the PR description and ask. ### Things AI assistants should NOT do - **Do not invent new top-level packages.** Add to existing packages instead. - **Do not rename public symbols** (router paths, dataclass field names, env vars, CLI flag names). External scripts depend on them. - **Do not add new infrastructure dependencies** (`boto3`, `redis`, `kafka`, `supabase`) at import time -- they are already optional and live in the `infra` extras. - **Do not write `monkeypatch` or `global` state in production code** -- use the repository pattern. - **Do not commit secrets, real prompts, or production paths.** The repo is scrubbed of these in CI. --- ## Common Tasks ### Add a new capability to an existing module 1. Create `studio//capabilities/.py` with a function signature `(request: CapabilityRequest) -> CapabilityResult`. 2. Register it in `studio//capabilities/__init__.py` (or the module's `registry.py`). 3. Add a test in `tests/unit/studio//test_.py`. ### Add a new HTTP endpoint 1. Identify which router it belongs to (or create one). One router is roughly one feature area, e.g. `studio_batch.py` for v4 batch operations. 2. Write the Pydantic request model at the top of the file. 3. Write the function -- keep it thin (<= 30 lines). 4. If it touches shared logic, look in `studio/reports.py` first. 5. Add an integration test under `tests/integration/test_.py`. ### Add a new bulk_* endpoint 1. Copy the closest sibling in `api/routers/studio_batch.py` (e.g. `bulk_swap_prompt` is the cleanest pattern to start from). 2. Write the Pydantic request model with `task_ids`, `by`, `include_locked`. 3. Define `def mutate(task): ...` returning a `BulkDecision`. Do the field change in place; the helper handles save + log + updated_at. 4. Call `bulk_apply_to_tasks(orch, req.task_ids, include_locked=..., mutate=mutate, now=now)`. 5. Map `result.updated / skipped_locked / skipped_other / errors` onto your response shape. Each endpoint has its own skip buckets -- see the section "Bulk endpoint pattern" above. 6. Add an integration test under `tests/integration/test_v4_e2e_http.py` using the `test_bulk_xxx_` naming convention. Cover at least: basic happy path, locked-skip, archived-skip (if relevant), and the domain-specific skip bucket (no_change, already_has, etc.). ### Run a single test file ```bash AIVS_SAAS_GATEWAY=0 .venv/bin/pytest tests/integration/test_v4_e2e_http.py -k "prompt_similarity_clusters" ``` ### Smoke-test the harness (closed loop) ```bash aivs harness doctor # confirm capability registry + 6 modules dry-run OK aivs harness run --topic "春日樱花" --material-dir ./material --json ``` ### Add a new LLM / media provider 1. Subclass the appropriate provider protocol in `studio//llm_provider.py`. 2. Add a unit test using the existing fake provider pattern. 3. Register it in the module's `__init__.py`. --- ## Operations Cheat-Sheet ```bash aivs maintenance doctor # full health check aivs maintenance validate-config # YAML schema check aivs maintenance retry task_xxx # retry a failed task aivs inbox archive # move processed inbox files aivs outbox list # inspect generated next-prompt outputs aivs harness doctor # capability registry + 6-module dry-run check aivs creation list-scenes # list builtin scenes aivs creation list-paradigms # list builtin paradigms (Hook / 3-act / B-roll ...) aivs creation bind-footage ... # bind Scene + Paradigm + StudioAsset into a Footage aivs creation stats # builtin / custom counts aivs harness explain # list registered capabilities aivs harness run --topic T # one-shot closed loop (understand -> remix) ``` Full ops playbook: `docs/05_operations_runbook.md`. --- ## Documentation Map | Topic | Doc | | --- | --- | | Layered architecture | `docs/00_architecture.md` | | Operations | `docs/01_operations.md` | | Provider guide | `docs/02_provider_guide.md` | | Desktop shell | `docs/03_desktop_shell.md` | | Storage backend | `docs/04_storage.md` | | Ops runbook | `docs/05_operations_runbook.md` | | v3 module loop | `docs/06_architecture_v3.md` | | LLM / enterprise | `docs/07_llm_and_enterprise.md` | | Real media + learning | `docs/08_real_media_and_learning.md` | | Architecture overview | `docs/09_architecture_overview.md` | | Capability registry | `docs/11_architecture_optimization.md` | | Atomic capabilities | `docs/12_module_splits.md` | | Data loop | `docs/13_data_loop.md` | | Deployment modes | `docs/14_deployment_modes.md` | | Runway / Adobe research | `docs/15_runway_adobe_research.md` | | v3.5-v3.7 implementation | `docs/16_v3_5_v3_7_implementation.md` | | v3.8 implementation | `docs/17_v3_8_implementation.md` | | v4 video-first UI | `docs/18_v4_video_studio_ui.md` | | Scene / paradigm / footage | `docs/19_creation_paradigm.md` | | Watermark remover | `docs/20_watermark_remover.md` | | Editing container runtime | `docs/21_editing_container_runtime.md` | | OCI / Podman topology | `docs/22_oci_topology_podman.md` | | Batch pipeline | `docs/23_batch_pipeline.md` | | Atomic local batch | `docs/23_osi_atomic_local_batch.md` | --- ## Contributing We welcome bug reports, docs improvements, and PRs. Please read `CONTRIBUTING.md` first; the `CODE_OF_CONDUCT.md` applies to all interactions. For security issues, follow `SECURITY.md` -- **do not** open a public issue. ## License [MIT](LICENSE).