# EnviroML **Repository Path**: zycbruceli/enviro-ml ## Basic Information - **Project Name**: EnviroML - **Description**: No description available - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-08-23 - **Last Updated**: 2026-09-08 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # EnviroML ## A leakage-aware spatiotemporal PM2.5 modelling and visualization platform EnviroML is an end-to-end environmental modelling software prototype for preparing air-quality observations, constructing temporal and spatial features, training multiple machine-learning models, and inspecting reproducible prediction results on an interactive GIS dashboard. The repository is prepared as the software companion for an **Environmental Modelling & Software (EMS)** submission. Its main methodological emphasis is a strict separation between model fitting data, evaluation data, and visualization-only meteorological context. The system is intended for research prototyping, method comparison, and reproducible inspection of station-level PM2.5 predictions. > This repository does not redistribute a research dataset. Users must provide data for which they have the necessary access and redistribution rights. ## Scope and main features - CSV-oriented dataset upload, profiling, field mapping, cleaning, and export. - Configurable temporal, lag, rolling-window, normalization, and spatial feature construction. - Chronological train/validation/test splitting with timestamp-level partition boundaries. - Multiple candidate models, including Random Forest, Gradient Boosting, and XGBoost, depending on the selected workflow. - FastAPI backend, Celery task execution, Redis broker/result backend, and a Vue 3 + TypeScript frontend. - Validation and test prediction tables, JSON summaries, GeoJSON map layers, reports, and provenance metadata. - GIS visualization of prediction points, heatmaps, administrative boundaries, spatial-analysis layers, and raw meteorological wind observations aligned to the selected evaluation partition. - English and Chinese user-interface modes. ## Methodological safeguards The following contracts are part of the current training and output pipeline. ### Partition-aware modelling 1. Timestamps are ordered before the chronological split. 2. All observations sharing a split timestamp are kept in the same partition. 3. The model is fitted on `train` only. 4. `validation` and `test` are transformed and predicted without refitting preprocessing state. 5. Validation is used for candidate-model selection and default model reporting; test remains a held-out evaluation partition. ### Fit/transform boundary The stateful `TrainOnlyPreprocessor` fits imputation, outlier, scaling, and categorical-processing state from the training split. Validation and test data use `transform` with that already-fitted state. The generated pipeline records split and preprocessing provenance in `pipeline_provenance.json`. ### Causal temporal features Lag and rolling features are grouped by station where a station identifier is available, ordered by time, and constructed from historical observations. Positive lags and causal rolling windows are used; future rows and backfilled target information are not used to create predictors. ### Target-derived feature protection Target-derived, AQI-like, target proxy, and metadata columns are not silently promoted to model features. The default generated workflow sets `allow_target_derived_features` to `false`. If a study intentionally enables a target-derived feature, that choice must be documented and scientifically justified. ### A single source of truth for evaluation maps For each partition, metrics, prediction tables, JSON payloads, and the corresponding map layer are derived from the same partition result table. In particular: - `validation_results.csv` is the canonical validation prediction table. - `test_results.csv` is the canonical test prediction table. - `validation_prediction_map.geojson` is derived from validation results. - `test_prediction_map.geojson` is derived from test results. - `full_predictions_diagnostics.csv` is diagnostic output only and is not a map source. ### Wind-layer isolation The wind layer is a visualization-only meteorological context. It reads wind observations from the current task's raw dataset, aligns them by normalized `station_id` and timestamp, and filters them to the selected validation or test time scope. Wind observations returned by the GIS endpoint are not fed back into feature engineering, model fitting, prediction, model selection, or metrics. ## System architecture ```text ┌──────────────────────────┐ │ Vue 3 / TypeScript UI │ │ MapLibre GIS dashboard │ └────────────┬─────────────┘ │ HTTP / JSON ┌────────────▼─────────────┐ │ FastAPI API │ │ datasets, jobs, analysis │ └───────┬─────────┬────────┘ │ │ SQLite metadata │ task dispatch │ ▼ │ Redis + Celery │ │ ┌───────▼─────────▼─────────┐ │ leakage-aware workflow │ │ preprocessing / training │ └────────────┬──────────────┘ ▼ backend/outputs/job_/ ``` ## Repository layout ```text EnviroML/ ├── backend/ │ ├── app/ # FastAPI routes, schemas, and services │ ├── worker/ # Celery application and task definitions │ ├── requirements.txt # Python dependencies │ ├── init_db.py # SQLite metadata database initialization │ ├── uploads/ # local uploaded datasets (not versioned) │ └── outputs/ # local job artifacts (not versioned) ├── frontend/ │ ├── src/ # Vue pages, components, GIS layers, and i18n UI │ ├── package.json # Node scripts and dependencies │ └── vite.config.ts # frontend dev server and API proxy └── README.md ``` ## Requirements - Windows, macOS, or Linux for the application services. Windows uses the `solo` Celery pool in the examples below. - Conda or another Python environment manager. - Python environment named `aisys` for the documented setup. The pinned scientific packages are intended for a modern Python 3 environment. - Node.js and npm for the frontend. - Redis 7 or a compatible Redis server. Docker Desktop is the recommended local option. - A dataset containing time, station identity, target, and coordinate fields suitable for the selected workflow. ## Installation ### 1. Create or activate the Python environment If the environment already exists: ```powershell conda activate aisys ``` For a new environment, create one first and then install the dependencies: ```powershell conda create -n aisys python=3.12 conda activate aisys ``` Install backend dependencies: ```powershell cd backend python -m pip install -r requirements.txt ``` Install frontend dependencies: ```powershell cd ..\frontend npm install ``` ### 2. Configure the backend The default local configuration uses SQLite for metadata and Redis on `localhost:6379`: ```text DATABASE_URL=sqlite:///./ai_modeling.db REDIS_URL=redis://localhost:6379/0 CELERY_BROKER_URL=redis://localhost:6379/0 CELERY_RESULT_BACKEND=redis://localhost:6379/0 ``` These defaults are already defined in `backend/app/core/config.py`. A local `backend/.env` may override them. `backend/.env.example` contains a PostgreSQL-oriented example and must be edited before use; do not copy it unchanged into a SQLite-only installation. Initialize the metadata database when setting up a fresh checkout: ```powershell cd backend python init_db.py ``` ## Running the system locally Open four terminals. Activate `aisys` in the terminals that run Python services. ### Redis Start Redis 7 or a compatible Redis service on `localhost:6379`. For a disposable local Docker instance: ```powershell docker run --name air_redis -p 6379:6379 -d redis:7-alpine ``` If `air_redis` already exists, start it with `docker start air_redis` instead. Configure `REDIS_URL`, `CELERY_BROKER_URL`, and `CELERY_RESULT_BACKEND` when Redis is hosted elsewhere. ### FastAPI backend ```powershell cd backend conda activate aisys python -m uvicorn app.main:app --reload --host 0.0.0.0 --port 8000 ``` The API is available at `http://localhost:8000`. The interactive API description is at `http://localhost:8000/docs`. ### Celery worker On Windows, use the following command exactly: ```powershell cd backend conda activate aisys python -m celery -A worker.celery_app:celery_app worker --loglevel=info --pool=solo ``` `--pool=solo` is the valid Celery option for this setup. `--pool-solo` is not a valid option. A healthy worker should eventually log a `ready` message and list the registered workflow tasks. ### Frontend ```powershell cd frontend npm run dev ``` Open `http://localhost:5173`. Vite proxies `/api` requests to `http://127.0.0.1:8000` by default. To use another backend address, set `VITE_PROXY_TARGET` before starting Vite. ## Data and workflow usage 1. Upload or register a task dataset in the data-management page. 2. Map the dataset's timestamp, station identity, target, longitude, latitude, and predictor fields. 3. Configure cleaning and feature engineering. Lag and rolling features should use a meaningful station identifier and a documented time resolution. 4. Configure the chronological train/validation/test proportions or the existing workflow split contract. 5. Select candidate models and submit the workflow. 6. Wait for the Celery worker to complete the task. 7. Inspect the partition-specific report and map. Use the `Validation` / `Test` selector to compare evaluation partitions without mixing their artifacts. Common target aliases such as `pm25` can be mapped to the canonical target field used by the dataset, for example `pm25_daily_mean`, through the field-mapping workflow. The exact required columns are dataset-dependent, but a typical station-level table contains: | Role | Example fields | | --- | --- | | Time | `date`, `timestamp` | | Station identity | `station_id`, `stationCode` | | Coordinates | `longitude`, `latitude` | | Target | `pm25`, `pm25_daily_mean` | | Meteorology | temperature, pressure, precipitation, wind speed/direction, or `u`/`v` components | | Other predictors | pollutant and environmental measurements that are available at prediction time | Do not treat a field as a legal predictor merely because it is present in the raw table. In particular, target copies, AQI calculations, future observations, post-outcome summaries, and fields derived from the full dataset require explicit review. ## Job artifacts Artifacts are written below `backend/outputs/job_/`. The exact set depends on the workflow, but a completed multi-model task normally contains the following groups: | Artifact | Meaning | | --- | --- | | `pipeline_provenance.json` | Split identity, feature/preprocessing hashes, target/time/station fields, and leakage-safety metadata | | `metrics.json` | Train, validation, and test metrics plus model-selection metadata | | `validation_results.csv` | Canonical validation sample-level predictions and errors | | `validation_predictions.json` | Validation metadata, metrics, and predictions in JSON form | | `validation_prediction_map.geojson` | Validation prediction points for GIS rendering | | `validation_prediction_map.json` | Validation timeline frames for the dashboard | | `test_results.csv` | Canonical held-out test sample-level predictions and errors | | `test_predictions.json` | Test metadata, metrics, and predictions in JSON form | | `test_prediction_map.geojson` | Test prediction points for GIS rendering | | `test_prediction_map.json` | Test timeline frames for the dashboard | | `full_predictions_diagnostics.csv` | Full diagnostic output; never use this as the source for an evaluation map | | `feature_importance_by_model.json` | Feature-importance summaries for available models | | `models/` | Serialized candidate and selected models where produced | | `charts/` and reports | Optional generated charts and HTML/Markdown reports | For every evaluation partition, the following should reconcile: ```text prediction table row count = partition prediction count in metadata = number of corresponding map records after coordinate exclusion is reported ``` Rows without valid coordinates are excluded from the map and counted in `excluded_map_point_count`; they are not silently removed from the evaluation table. ## Metrics and evaluation semantics The result tables expose `actual`, `predicted`, `residual`, and `absolute_error` fields. Here `residual = actual - predicted`, so a positive residual indicates under-prediction and a negative residual indicates over-prediction. The standard metrics are calculated on the selected partition: - `R²`: coefficient of determination. - `RMSE`: root mean squared error. - `MAE`: mean absolute error. Validation metrics are used for candidate-model selection in the generated multi-model workflow. Test metrics are retained for final held-out assessment. The UI can switch the report and map between validation and test data; the selected partition must always be visible in the report metadata. For an independent artifact check, recompute metrics directly from the relevant CSV: ```python import numpy as np import pandas as pd from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score results = pd.read_csv("backend/outputs/job_/validation_results.csv") y_true = results["actual"] y_pred = results["predicted"] print("R2 ", r2_score(y_true, y_pred)) print("RMSE", np.sqrt(mean_squared_error(y_true, y_pred))) print("MAE ", mean_absolute_error(y_true, y_pred)) ``` ## Wind visualization semantics The wind endpoint is exposed through the analysis API as: ```text GET /api/v1/analysis/jobs/{job_id}/layers/wind ``` Wind data is read from the selected job's source dataset, not from another job or the most recently uploaded file. The service supports explicit aliases for speed/direction fields and `u`/`v` component fields. If only direction is available, the layer can still be displayed and reports that speed is unavailable. If only speed is available, the direction-dependent arrow rendering is unavailable. The service aligns wind observations to the current evaluation partition using normalized station identity and timestamps. It then filters the frontend layer to the current timeline timestamp. It does not silently perform an unsupported nearest-timestamp join. The direction convention must be documented for each dataset. Meteorological direction commonly means the direction **from which** the wind comes, while a map vector usually points **towards** where the wind moves. Any conversion used by the renderer must be explicit in the wind-layer metadata and code comments. ## Quality checks Build the frontend before deployment: ```powershell cd frontend npm run build ``` The submitted source tree intentionally excludes local repair scripts, migration helpers, serialized models, generated outputs, and test-only fixtures. For a manuscript submission, retain the test reports and experiment configurations as supplementary reproducibility material if they are not distributed with the software package. ## Troubleshooting ### The worker remains in the startup phase - Confirm Redis responds on `localhost:6379`. - Confirm only one worker is consuming the development queue. - Run the Windows command with `--pool=solo`. - Inspect the worker log for a final `ready` message before submitting a task. ### A task stays queued - Check that the API and worker use the same `CELERY_BROKER_URL`. - Check that both processes are running from `backend/` or use the same environment configuration. - Confirm the worker lists `worker.tasks.automl_tasks.execute_automl_workflow` and `worker.tasks.training_tasks.train_model_task`. ### The frontend cannot reach the API - Check `http://localhost:8000/docs` first. - Check the Vite proxy target and restart Vite after changing it. - Verify that port `8000` is not occupied by an unrelated process. ### A map is empty while the table has rows - Confirm that the selected partition matches the map endpoint. - Check `validation_results.csv` or `test_results.csv` rather than `full_predictions_diagnostics.csv`. - Check `excluded_map_point_count` and coordinate validity. - Check the job's `pipeline_provenance.json` and source dataset identity. ## Reproducibility and reporting checklist When using EnviroML in a paper or benchmark, report at least: 1. Dataset identity, licensing/access conditions, temporal coverage, and station coverage. 2. Target, timestamp, station, coordinate, and feature-field mappings. 3. Train/validation/test split strategy and date ranges. 4. Cleaning, lag, rolling, normalization, spatial-feature, and missing-data settings. 5. Candidate models, hyperparameters, random seeds, and selection metric scope. 6. Train, validation, and held-out test `R²`, `RMSE`, and `MAE`. 7. The relevant `pipeline_provenance.json` and artifact counts. 8. Whether wind and other raw meteorological layers were used only for visualization or also as legal model predictors. 9. Any missing-coordinate exclusions, station exclusions, or unsupported time alignments. ## Limitations - This repository is a research-oriented prototype rather than a production distributed platform. - Results depend on the quality, temporal resolution, station coverage, and field semantics of the supplied dataset. - A chronological split reduces temporal leakage but does not by itself establish causal validity or eliminate spatial representativeness bias. - A model can have acceptable aggregate metrics while showing persistent station-level residual bias. Residuals should therefore be inspected by station and time, not only summarized globally. - External basemap tiles require network access and must be used with the provider's attribution and terms. - No universal default can determine whether a domain-specific field is available at prediction time; field provenance remains a study-level responsibility. ## Citation and license If you use EnviroML in a publication, please cite the associated EMS manuscript once its bibliographic information is available. Until a project license and citation record are formally added to this repository, contact the project authors before redistributing the software or bundled artifacts. ```text EnviroML: A leakage-aware spatiotemporal PM2.5 modelling and visualization platform. Software companion for an Environmental Modelling & Software submission. ```