schema.py · config.py · definitions.py

The pipeline's contract, configuration, and wiring layer — not a data flow, a dependency graph

Click any block for details

schema.py
Data contracts — no Dagster, no compute
config.py
CarbonPipelineConfig, StorageConfig
config.py imports DEFAULT_VALIDATION_LEVEL from schema.py →
assets/cpu/*.py + assets/gpu/*.py
each asset takes CarbonPipelineConfig and reads schema.py constants directly
resources/carbon.py
CarbonModelResource — model + tokenizer
resources/hf_client.py
create_huggingface_resource()
definitions.py — imports assets + resources, wires everything together
Asset groupings
CPU_ASSETS, GPU_PIPELINE_ASSETS
Job definitions
4 jobs, each an AssetSelection
defs = dg.Definitions(assets, jobs, resources)
the single object Dagster actually loads
used by downstream analysis & catalog sync — not registered in dg.Definitions() ↓
resources/*.py — analytical & catalog access layer (standalone)
faceberg.py
PipelineCatalog, PIPELINE_TABLES lineage registry
duckdb.py
DuckDB + Iceberg query surface
clickhouse.py
ClickHouseResource — local analytical SQL
lancedb.py
LanceDBResource — vector storage/search
Detail
Click any block above to see how it works.
Contracts Runtime config Asset modules Resource modules Asset groupings Jobs Definitions object