# Attribute dataset construction **Repository Path**: wang-kun_bo/attribute-dataset-construction ## Basic Information - **Project Name**: Attribute dataset construction - **Description**: No description available - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 1 - **Forks**: 0 - **Created**: 2026-05-13 - **Last Updated**: 2026-06-11 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # Region-Level Discriminative Cue Dataset Builder This repository builds a region-level dataset with: - Prompt A: main annotation - Prompt B: visual consistency validation - Prompt C: cue canonicalization The code supports separate teacher roles: - `vision_teacher` for Prompt A / Prompt B - `text_teacher` for Prompt C The user prepares models, APIs, and datasets on the server. This repository only provides code, configs, prompts, and runnable pipelines. ## What It Produces The pipeline preserves three cue layers: - `raw_cues` - `validated_cues` - `canonical_cues` It also preserves: - teacher raw responses - parsed responses - failure records - provenance metadata - timestamps ## Project Layout ```text src/cue_dataset_builder/ cli.py config/ converters/ data/ exporter/ loaders/ pipeline/ prompts/ qa/ schemas/ teacher/ utils/ vocab/ configs/ examples/ scripts/ docs/ ``` ## Environment ```bash conda create -n cue-dataset-builder python=3.10 -y conda activate cue-dataset-builder pip install -e . ``` ## Prompt Assignment Default teacher split: - Prompt A / B: vision-language model - Prompt C: text model The code does not assume all stages use the same model. ## Main Schemas `RegionInputSample` - `id` - `image` - `split` - `region_source` - `bbox` - optional `mask_path` / `polygon` - `metadata` `AnnotationAResult` - `reject` - `reject_reason` - `coarse_category` - `category` - `raw_cues` `ValidationBResult` - `reject` - `reject_reason` - `category_valid` - `validated_cues` - `invalid_cues` `CanonicalCResult` - `canonical_cues` - `cue_types` `FinalTrainingSample` - `id` - `image` - `bbox` - optional `mask` - `coarse_category` - `category` - `raw_cues` - `validated_cues` - `canonical_cues` - `cue_types` - `teacher_model` - `prompt_version` - `timestamps` - `provenance` ## CLI ```bash cue-dataset validate-config --config [--check-files] cue-dataset convert-coco --annotations --image-dir --output cue-dataset build-regions --config [--limit 100] cue-dataset annotate-a --config [--limit 100] [--dry-run] cue-dataset validate-b --config [--limit 100] [--dry-run] cue-dataset canonicalize-c --config [--limit 100] [--dry-run] cue-dataset build-vocab --config cue-dataset apply-vocab --config cue-dataset build-paper-attribute-pool --config cue-dataset select-concise-cues --config [--top-k 256] cue-dataset select-task-guided-cues --config [--top-k 256] cue-dataset sample-qa --config cue-dataset render-qa-html --config [--sample-size 100] [--shuffle] cue-dataset report-stats --config cue-dataset merge-human-fixes --config --fixes cue-dataset export-dataset --config cue-dataset evaluate-cue-guess --config [--cue-source canonical] [--clip-model openai/clip-vit-base-patch32] For a medium-sized training test set, use [configs/pipeline.coco.openai_compatible.5k.yaml](./configs/pipeline.coco.openai_compatible.5k.yaml). The export folder will also include `exports/DATASET_FORMAT.md` describing the saved file formats. For a broader-category, more disentanglement-oriented run, use [configs/pipeline.coco.openai_compatible.disentangle.yaml](./configs/pipeline.coco.openai_compatible.disentangle.yaml). cue-dataset preview-stage --config --stage [--limit 3] cue-dataset inspect-sample --config --id cue-dataset run-all --config [--limit 100] [--dry-run] ## Concise Cue Filtering If you already exported a dataset and want a cleaned copy without touching the original `final_training_samples.jsonl`, you can build a smaller global cue vocabulary and write a separate filtered file. The lightweight version ranks canonical cues using a combination of: - global usage count - category coverage - cross-category entropy - penalty for strongly category-like cue phrases Example: ```bash cue-dataset select-concise-cues \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --top-k 256 \ --min-global-count 2 ``` Default outputs: - `exports/final_training_samples.concise.jsonl` - `vocab/selected_concise_cues.jsonl` - `vocab/concise_cue_report.json` This command does not overwrite the original exported dataset. ## Task-Guided Cue Selection If you want to follow the paper more closely from the attribute-pool construction step, you can first build a class-driven attribute pool from category names using instance and batch prompting. Example: ```bash cue-dataset build-paper-attribute-pool \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --teacher-role text \ --prompt-version paper_v1 \ --group-size 8 ``` Default outputs: - `vocab/paper_attribute_pool/paper_v1/instance.jsonl` - `vocab/paper_attribute_pool/paper_v1/batch.jsonl` - `vocab/paper_attribute_pool/paper_v1/merged.jsonl` - `vocab/paper_attribute_pool/paper_v1/report.json` You can then feed that merged attribute pool into task-guided selection: ```bash cue-dataset select-task-guided-cues \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --cue-pool-input artifacts/coco_val2017_openai_compatible_disentangle/vocab/paper_attribute_pool/paper_v1/merged.jsonl \ --clip-model ./models/clip/ViT-B-16.pt \ --clip-device cuda:0 \ --top-k 256 \ --min-global-count 2 \ --mahalanobis-lambda 1e-2 ``` If you want the downstream selection step to more closely follow the paper, add `--paper-strict`. This disables the extra pool prefilters and retrains a fresh linear probe after nearest-neighbor retrieval. For a more paper-like version, you can use paired image/cue embedding similarities together with category labels to learn a small cue dictionary and then retrieve the nearest neighbor cues from the global cue pool. By default, this command: 1. crops each region from the exported dataset using its bbox 2. encodes region crops with a VLM image encoder 3. encodes the global cue pool with the matching text encoder 4. learns a task-guided dictionary with classification supervision 5. regularizes the dictionary toward the cue text embedding distribution 6. retrieves the nearest neighbor cues from the original global cue pool 7. writes a new filtered dataset without touching the original export Example: ```bash cue-dataset select-task-guided-cues \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --clip-model ./models/clip/ViT-B-16.pt \ --clip-device cuda:0 \ --top-k 256 \ --min-global-count 2 \ --mahalanobis-lambda 1e-2 ``` If you already have another dual-tower feature space aligned across image regions and cue texts (for example, WeDetectUni image features plus WeDetectUni attribute-text features), you can swap either tower or both towers through sidecar JSONL files. The only requirement is that the image and cue sidecars must have the same embedding dimension and live in the same semantic space. Expected sidecar formats: - image sidecar JSONL: `{"id": "...", "feature": [...]}` - cue sidecar JSONL: `{"cue": "...", "feature": [...]}` Example with WeDetectUni sidecars for both towers: ```bash cue-dataset select-task-guided-cues \ --config /space0/liangc/others/wangkb/attribute-dataset-construction/configs/pipeline.coco.openai_compatible.disentangle.yaml \ --input /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/exports/final_training_samples.jsonl \ --cue-pool-input /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/paper_attribute_pool/paper_v1/merged.jsonl \ --image-feature-source sidecar_jsonl \ --image-feature-sidecar-jsonl /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/features/wedetectuni_features.jsonl \ --image-sidecar-id-field id \ --image-feature-field feature \ --cue-feature-source sidecar_jsonl \ --cue-feature-sidecar-jsonl /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/features/wedetectuni_attribute_features.paper_v1.jsonl \ --cue-sidecar-key-field cue \ --cue-feature-field feature \ --top-k 256 \ --paper-strict \ --mahalanobis-lambda 1e-2 \ --report-output /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/task_guided_cue_report.wedetectuni.paper_strict.json \ --selected-output /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/selected_task_guided_cues.wedetectuni.paper_strict.jsonl ``` If you want a stronger nonlinear readout for diagnostics, you can replace the default linear probe with a lightweight MLP probe: ```bash cue-dataset select-task-guided-cues \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --clip-model ./models/clip/ViT-B-16.pt \ --clip-device cuda:0 \ --top-k 256 \ --paper-strict \ --probe-type mlp \ --probe-hidden-dim 512 \ --probe-dropout 0.1 ``` Default outputs: - `exports/final_training_samples.task_guided.jsonl` - `vocab/selected_task_guided_cues.jsonl` - `vocab/task_guided_cue_report.json` The selected cue JSONL now also includes simple probe-based importance diagnostics for each selected attribute: - `dictionary_importance_score` / `dictionary_importance_rank` - `retrieved_importance_score` / `retrieved_importance_rank` These scores are derived from the classifier weights over the selected top-k dimensions. For the default linear probe they are exact input-weight norms; for the optional MLP probe they are an approximation based on the first linear layer. ## Task-guided Sweep If you want to compare different `K` values and Mahalanobis regularization strengths in one batch, you can run `sweep-task-guided-cues`. It saves one report + selected JSONL per run and also writes: - `summary.json` - `summary.csv` - `attribute_summary.csv` - `attribute_details.csv` - `summary.html` The HTML page highlights both per-run metrics and the attributes that are repeatedly selected or have high retrieved-importance scores. Example: ```bash cue-dataset sweep-task-guided-cues \ --config /space0/liangc/others/wangkb/attribute-dataset-construction/configs/pipeline.coco.openai_compatible.disentangle.yaml \ --input /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/exports/final_training_samples.jsonl \ --cue-pool-input /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/paper_attribute_pool/paper_v1/merged.jsonl \ --image-feature-source sidecar_jsonl \ --image-feature-sidecar-jsonl /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/features/wedetectuni_features.jsonl \ --cue-feature-source sidecar_jsonl \ --cue-feature-sidecar-jsonl /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/features/wedetectuni_attribute_features.paper_v1.jsonl \ --top-k-values 16 32 64 128 256 512 \ --mahalanobis-lambda-values 0 1e-3 1e-2 1e-1 \ --paper-strict \ --sweep-output-dir /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/sweeps/wedetectuni_paper_v1 ``` ## Task-guided Distribution View If you want to visualize the cue/text feature distribution before and after selection for one specific task-guided run, you can render a lightweight HTML scatter view. It projects the full cue pool into 2D with PCA, then highlights the selected top-k cues in the same projection. Example with a WeDetectUni canonical run: ```bash cue-dataset render-task-guided-distribution-html \ --config /space0/liangc/others/wangkb/attribute-dataset-construction/configs/pipeline.coco.openai_compatible.disentangle.yaml \ --input /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/exports/final_training_samples.jsonl \ --selected-input /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/selected_task_guided_cues.wedetectuni.canonical_strict.jsonl \ --output-html /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/selected_task_guided_cues.wedetectuni.canonical_strict.distribution.html \ --cue-feature-source sidecar_jsonl \ --cue-feature-sidecar-jsonl /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/features/wedetectuni_attribute_features.canonical.jsonl \ --cue-sidecar-key-field cue \ --cue-feature-field feature \ --paper-strict ``` If you already have several selected JSONL files for different lambda values, you can overlay them on the same PCA projection and use one color per run: ```bash cue-dataset render-task-guided-multi-distribution-html \ --config /space0/liangc/others/wangkb/attribute-dataset-construction/configs/pipeline.coco.openai_compatible.disentangle.yaml \ --input /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/exports/final_training_samples.jsonl \ --selected-inputs \ /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/vis_k128_lambda/selected.canonical.k128.lambda0.jsonl \ /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/vis_k128_lambda/selected.canonical.k128.lambda0p01.jsonl \ /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/vis_k128_lambda/selected.canonical.k128.lambda1.jsonl \ --labels lambda0 lambda0p01 lambda1 \ --output-html /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/vocab/vis_k128_lambda/distribution.canonical.k128.multi_lambda.html \ --cue-feature-source sidecar_jsonl \ --cue-feature-sidecar-jsonl /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/features/wedetectuni_attribute_features.canonical.jsonl \ --cue-sidecar-key-field cue \ --cue-feature-field feature \ --paper-strict ``` This is closer to the paper than the lightweight concise filter, but it is still an adaptation for the current project. The main differences are: - the paper is formulated for image-level recognition datasets, while this project uses region crops from bbox samples - the train/validation split is created inside the exported dataset - hyperparameters are practical defaults for the current pipeline rather than a full paper-style sweep ## Feature Top-k Evaluation If you want to compare raw feature quality against a task-guided top-k bottleneck without forcing the features to align with text attributes, you can run `evaluate-feature-topk`. This command always reports two layers: 1. `raw_feature_probe_metrics`: a probe trained directly on the input features 2. `topk_dictionary_probe_metrics`: a probe trained on `x @ D^T`, where `D` is a learned top-k dictionary Example with CLIP region features: ```bash cue-dataset evaluate-feature-topk \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --feature-source clip_region \ --clip-model ./models/clip/ViT-B-16.pt \ --clip-device cuda:0 \ --train-device cuda:0 \ --top-k 256 \ --report-output artifacts/coco_val2017_openai_compatible_disentangle/vocab/feature_topk_report.clip_region.json ``` Example with precomputed features already stored inside each JSONL record: ```bash cue-dataset evaluate-feature-topk \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --feature-source record_field \ --feature-field wedetectuni_feature \ --train-device cuda:0 \ --top-k 256 \ --report-output artifacts/coco_val2017_openai_compatible_disentangle/vocab/feature_topk_report.wedetectuni.json ``` If your feature field is nested, use dotted paths such as: - `wedetectuni.feature` - `features.wedetectuni.vector` Example with a sidecar JSONL file aligned by sample `id`: ```bash cue-dataset evaluate-feature-topk \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --feature-source sidecar_jsonl \ --feature-sidecar-jsonl /space0/liangc/others/wangkb/attribute-dataset-construction/artifacts/coco_val2017_openai_compatible_disentangle/features/wedetectuni_features.jsonl \ --sidecar-id-field id \ --feature-field feature \ --train-device cuda:0 \ --top-k 256 \ --report-output artifacts/coco_val2017_openai_compatible_disentangle/vocab/feature_topk_report.wedetectuni.json ``` ## Cue Guess Evaluation The `evaluate-cue-guess` experiment does not measure closed-set classification accuracy. Instead, it tests whether a cue list can naturally induce a semantically close category guess. Pipeline: 1. Feed `cue_list` into a text LLM using an open-set guessing prompt. 2. Parse `predicted_category` and related reasoning fields. 3. Encode `predicted_category` and ground-truth `category` with CLIP text prompts: - `a photo of a {predicted_category}` - `a photo of a {ground_truth_category}` 4. Use cosine similarity between the two CLIP text embeddings as the main metric. Example: ```bash cue-dataset evaluate-cue-guess --config configs/pipeline.coco.openai_compatible.disentangle.yaml --cue-source canonical ``` Cue-guess prompt variants are selectable for prompt comparison experiments: - `open_set_v1`: the original open-set semantic guessing prompt - `conservative_v2`: a more conservative semantic evidence evaluator prompt Examples: ```bash cue-dataset evaluate-cue-guess \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --cue-source canonical \ --prompt-version open_set_v1 cue-dataset evaluate-cue-guess \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --cue-source canonical \ --prompt-version conservative_v2 ``` The output and cache paths are separated automatically by both teacher model and prompt version, for example: - `eval/cue_guess/qwen3_8b/open_set_v1/results.jsonl` - `eval/cue_guess/qwen3_8b/conservative_v2/results.jsonl` - `cache/cue_guess/qwen3_8b/open_set_v1/*.json` - `cache/cue_guess/qwen3_8b/conservative_v2/*.json` To visually inspect the lowest-scoring cue-guess samples, render an HTML review page: ```bash cue-dataset render-cue-guess-html \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --teacher-role text \ --prompt-version conservative_v2 \ --limit 200 ``` This generates a page sorted by low `clip_similarity` first, with image preview, bbox, ground-truth category, predicted category, cues, and reasoning. You can also inspect a middle slice instead of only the lowest-scoring block by combining `--offset` and `--limit`. For example, to render items ranked 1001-1200 after sorting by low score first: ```bash cue-dataset render-cue-guess-html \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --teacher-role text \ --prompt-version conservative_v2 \ --offset 1000 \ --limit 200 ``` By default, cue guessing uses `text_teacher`. You can also run the same cue-guess prompt with `vision_teacher` for a parallel comparison experiment: ```bash cue-dataset evaluate-cue-guess \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --cue-source canonical \ --teacher-role text cue-dataset evaluate-cue-guess \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --cue-source canonical \ --teacher-role vision ``` The outputs are separated automatically by teacher model name, for example: - `eval/cue_guess/qwen3_8b/results.jsonl` - `eval/cue_guess/qwen3_5_9b/results.jsonl` The CLIP scorer supports two loading modes: - Hugging Face / `transformers` model name or local directory, for example: - `--clip-model openai/clip-vit-base-patch32` - Local OpenAI CLIP `.pt` checkpoint, for example: - `--clip-model ./models/clip/ViT-B-16.pt` If you already have an OpenAI CLIP checkpoint on the server, you do not need to clone the CLIP repository. You only need the Python package to be importable in the evaluation environment. Recommended server-side layout: ```bash cd /space0/liangc/others/wangkb/attribute-dataset-construction mkdir -p models/clip ln -s /space0/liangc/.cache/clip/ViT-B-16.pt models/clip/ViT-B-16.pt ``` Then run: ```bash cue-dataset evaluate-cue-guess \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --cue-source canonical \ --clip-model ./models/clip/ViT-B-16.pt ``` ## LLM-as-Judge Evaluation The repository also supports a post-hoc `LLM-as-Judge` semantic matching pass on top of existing cue-guess results. This does **not** rerun cue guessing. It only reads existing `results.jsonl` files, asks a judge model to score the semantic match between `ground_truth_category` and `predicted_category`, and writes new judged outputs. The main use case is comparing prompt/model variants that have already finished running, for example: - `qwen3_8b/open_set_v1/results.jsonl` - `qwen3_8b/conservative_v2/results.jsonl` - `qwen3_5_9b/open_set_v1/results.jsonl` - `qwen3_5_9b/conservative_v2/results.jsonl` Single-file usage: ```bash python scripts/evaluate_llm_judge.py \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --input_file artifacts/coco_val2017_openai_compatible_disentangle/eval/cue_guess/qwen3_8b/open_set_v1/results.jsonl ``` Directory usage: ```bash python scripts/evaluate_llm_judge.py \ --config configs/pipeline.coco.openai_compatible.disentangle.yaml \ --input_dir artifacts/coco_val2017_openai_compatible_disentangle/eval/cue_guess \ --output_dir artifacts/coco_val2017_openai_compatible_disentangle/eval/llm_judge \ --judge-teacher-role text \ --temperature 0 \ --top-p 1 ``` This produces: - per-result judged files: `results.judge.jsonl` - per-result summaries: `summary.json` - per-result `bad_cases.jsonl` - directory-level `comparison_summary.json` - directory-level `comparison_summary.csv` Each judged record keeps the original cue-guess fields and appends: - `judge_score` - `judge_match_level` - `judge_is_synonym` - `judge_is_hypernym` - `judge_is_hyponym` - `judge_is_sibling_or_related` - `judge_is_invalid_prediction` - `judge_reason` - `judge_parse_failed` By default, the judge reuses `text_teacher`. You can also point it at a different teacher role or override the served model name / API base from the command line without touching the original cue-guess outputs. Default outputs: - `eval/cue_guess/results.jsonl` - `eval/cue_guess/summary.json` - `eval/cue_guess/bad_cases.jsonl` ``` ## COCO Example Use the built-in OpenAI-compatible example config: - [configs/pipeline.coco.openai_compatible.example.yaml](configs/pipeline.coco.openai_compatible.example.yaml) If your server has: - `/datasets/coco/annotations/instances_val2017.json` - `/datasets/coco/val2017` then the default COCO example config is already pointed at those paths. ### Step 1: validate config ```bash cue-dataset validate-config --config configs/pipeline.coco.openai_compatible.example.yaml --check-files ``` Before running Prompt A / B / C, start the two local teachers on the server: ```bash bash scripts/start_qwen_teachers.sh curl http://127.0.0.1:8000/v1/models curl http://127.0.0.1:8001/v1/models ``` The default server-side setup assumes: - `models/Qwen3.5-9B` for the vision teacher on port `8000` - `models/Qwen3-8B` for the text teacher on port `8001` - `OPENAI_API_KEY=EMPTY` if your local OpenAI-compatible service does not require authentication - `chat_template_kwargs.enable_thinking: false` for both teachers - `temperature=0.7`, `top_p=0.8`, `top_k=20` to match the official non-thinking guidance from the Qwen model cards For tighter GPU memory, use smaller startup defaults: ```bash VISION_GPU_MEMORY_UTILIZATION=0.55 \ TEXT_GPU_MEMORY_UTILIZATION=0.50 \ VISION_MAX_MODEL_LEN=4096 \ TEXT_MAX_MODEL_LEN=4096 \ bash scripts/start_qwen_teachers.sh ``` If you cannot keep both services resident at the same time, run them sequentially: ```bash START_TEXT_TEACHER=0 bash scripts/start_qwen_teachers.sh cue-dataset annotate-a --config configs/pipeline.coco.openai_compatible.example.yaml --limit 20 cue-dataset validate-b --config configs/pipeline.coco.openai_compatible.example.yaml --limit 20 bash scripts/stop_qwen_teachers.sh START_VISION_TEACHER=0 START_TEXT_TEACHER=1 bash scripts/start_qwen_teachers.sh cue-dataset canonicalize-c --config configs/pipeline.coco.openai_compatible.example.yaml --limit 20 cue-dataset export-dataset --config configs/pipeline.coco.openai_compatible.example.yaml ``` ### Step 2: build region input samples ```bash cue-dataset build-regions --config configs/pipeline.coco.openai_compatible.example.yaml --limit 100 ``` ### Step 3: run Prompt A / B / C ```bash cue-dataset annotate-a --config configs/pipeline.coco.openai_compatible.example.yaml --limit 20 cue-dataset validate-b --config configs/pipeline.coco.openai_compatible.example.yaml --limit 20 cue-dataset canonicalize-c --config configs/pipeline.coco.openai_compatible.example.yaml --limit 20 ``` ### Step 4: export training dataset ```bash cue-dataset export-dataset --config configs/pipeline.coco.openai_compatible.example.yaml ``` ### Step 5: inspect the outputs ```bash cue-dataset preview-stage --config configs/pipeline.coco.openai_compatible.example.yaml --stage annotate_a --limit 2 cue-dataset preview-stage --config configs/pipeline.coco.openai_compatible.example.yaml --stage canonicalize_c --limit 2 cue-dataset preview-stage --config configs/pipeline.coco.openai_compatible.example.yaml --stage final_training --limit 2 cue-dataset inspect-sample --config configs/pipeline.coco.openai_compatible.example.yaml --id coco:45596:124742 ``` `preview-stage` is useful when you want a quick glance at one stage output. `inspect-sample` is useful when you want to trace the same sample across build-regions, Prompt A, Prompt B, Prompt C, and final export. ### One-command dry run ```bash cue-dataset run-all --config configs/pipeline.coco.example.yaml --limit 50 --dry-run ``` `pipeline.coco.example.yaml` uses mock teachers for cheap structure validation. ## Prompt Files Prompt templates are packaged in: - `src/cue_dataset_builder/prompts/catalog.py` They include: - Prompt A: main annotation - Prompt B: consistency validation - Prompt C: cue canonicalization ## Outputs Under `output.root_dir`, the pipeline writes: - `regions/*.jsonl` - `stages/*.jsonl` - `stages/*.failures.jsonl` - `cache/*` - `exports/*` - `vocab/*` - `qa/*` - `logs/*` ## Notes - Models are not downloaded by this repository. - Datasets are not downloaded by this repository. - All paths are passed through configs or CLI. - Bbox is the current primary path. - Mask fields are preserved for future extension.