Cortical sulci — the folds of the brain surface — vary in shape across individuals and are linked to cognitive function, development, and neurological conditions. The Champollion pipeline turns T1 MRI scans into compact, comparable representations of sulcal morphology using self-supervised contrastive learning. It processes MRIs through the BrainVISA/Morphologist toolchain to extract sulcal graphs, uses cortical_tiles to crop standardized 3D patches around 28 sulcal regions per hemisphere, and then runs pre-trained Champollion encoders to produce 32-dimensional embeddings per region. These embeddings can be projected onto pre-trained UMAP reference maps for visualization and compared across cohorts. The pipeline is designed for researchers who want to apply Champollion to their own neuroimaging datasets without retraining.
Project website: https://www.neurospin.fr/champollion_pipeline
Try it online: A live demo is available on Hugging Face Spaces. It runs on 2 CPU cores and is suited for quick testing with a single subject. For batch processing or production use, install this pipeline locally where it can leverage all available CPUs and GPUs.
The pipeline runs in 6 steps. Here is the minimal happy path — replace placeholders '/data/myproject' with your actual paths:
# 1. Generate Morphologist sulcal graphs (requires BrainVISA — skip if you already have .arg files)
# # This takes as inputs the list of T1 MRI (here: sub-001.nii.gz sub-002.nii.gz)
# # and outputs the morphologist graphs into the folder /data/myproject/derivatives/morphologist-6.0/subjects
morphologist-cli sub-001.nii.gz sub-002.nii.gz /data/myproject/ \
-- --of morphologist-auto-nonoverlap-1.0 --if morphologist-auto-nonoverlap-1.0
# 2. Extract sulcal regions
# # Input: morphologist subjects directory (read-only, never written to)
# # Output: /data/myproject/derivatives/cortical_tiles-2026/crops/2mm/
pixi run champollion-cortical-tiles \
/data/myproject/derivatives/morphologist-6.0/subjects \ # morphologist output dir
/data/myproject/derivatives/ # derivatives root (crops land here)
# 3. Generate Champollion configuration
# # Links your crop paths to the Champollion model config
pixi run champollion-config \
/data/myproject/derivatives/cortical_tiles-2026/crops/2mm \ # path to 2mm crops (step 2 output)
--dataset myproject \ # dataset name used throughout the pipeline
--output /data/myproject/derivatives/champollion_V1/configs # where to write YAML config files
# 4. Generate embeddings (downloads pre-trained models from Hugging Face)
pixi run champollion-embeddings \
neurospin/Champollion_V1 \ # model source: HF repo ID, local path, or archive
local \ # localization preset (matches local.yaml written in step 3)
myproject \ # dataset name (must match --dataset from step 3)
my_run \ # label for this run (used to name output CSV files)
--embeddings_only \ # skip classifier training, compute embeddings only
--config_path /data/myproject/derivatives/champollion_V1/configs/dataset/myproject # step 3 output
# 5. Combine embeddings
# # Collects 56 per-region CSVs into a single output directory
pixi run champollion-combine \
--path_models /data/myproject/derivatives/champollion_V1/models_cache/Champollion_V1/ \ # models cache dir (step 4)
--embeddings_subpath my_run_random_embeddings/full_embeddings.csv \ # {run_label}_{split}_embeddings/full_embeddings.csv
--output_path /data/myproject/derivatives/champollion_V1/embeddings/ # destination for combined CSVs
# 6. Generate visualization snapshots
pixi run champollion-snapshots \
--embeddings_dir /data/myproject/derivatives/champollion_V1/embeddings/ \ # step 5 output
--reference_data_dir reference_data/ \ # pre-trained UMAP models (in repo)
--output_dir /data/myproject/derivatives/champollion_V1/snapshots/ # where to write imagesThe sections below explain each step in detail.
- Pixi package manager
- Git
- BrainVISA / Morphologist — required for step 2 only (sulcal graph extraction). If you already have Morphologist
.arggraphs, you can skip step 2 and start from step 3.
Clone the repository and install all dependencies:
mkdir Champollion && cd Champollion
git clone https://github.com/neurospin/champollion_pipeline.git
cd champollion_pipeline
pixi run install-allinstall-all initializes the git submodules (champollion_V1 and cortical_tiles), installs them in editable mode, clones champollion_utils, and creates the data/ directory.
To enter the managed environment interactively:
pixi shellpixi run uninstall # Remove installed packages and cloned repos
pixi run uninstall-all # Also remove Pixi-managed dependencies
⚠️ uninstall-allremoves thedata/folder as well. Back up any data stored there first.
pixi run updateIf you hit issues (submodule URL mismatch, missing hatchling, stale pixi tasks, etc.),
see migration_manual.md for symptoms and fixes.
Skip this step if you already have Morphologist sulcal graph files (
.argformat) for your subjects.
Morphologist is a tool from the BrainVISA neuroimaging suite that segments T1 MRI images and produces a graph representation of each subject's sulcal folds. It runs in its own BrainVISA environment, not through Pixi.
morphologist-cli sub-001.nii.gz sub-002.nii.gz /data/myproject/ \
-- --of morphologist-auto-nonoverlap-1.0 --if morphologist-auto-nonoverlap-1.0This writes one folder per subject under /data/myproject/derivatives/morphologist-6.0/subjects/.
For large cohorts, soma-workflow (BrainVISA's HPC job scheduler) can parallelize graph generation across CPU cores or cluster nodes. This is entirely optional — serial processing works fine for small cohorts.
First configure soma-workflow:
soma_workflow_gui # Set max CPUs under "Computing resources"Then add --swf:
morphologist-cli sub-001.nii.gz ... /data/myproject/ \
-- --of morphologist-auto-nonoverlap-1.0 --if morphologist-auto-nonoverlap-1.0 --swfExtract standardized 3D patches around 28 sulcal regions per hemisphere using cortical_tiles:
pixi run champollion-cortical-tiles \
/data/myproject/derivatives/morphologist-6.0/subjects \ # morphologist subjects dir (read-only)
/data/myproject/derivatives/ # derivatives root; crops land at {root}/cortical_tiles-2026/crops/2mm/- First argument — directory containing one folder per subject (Morphologist's
subjects/output). This path may be read-only (e.g. a shared NFS database); the script never writes to it. - Second argument — derivatives parent directory. Crops are written to
{output}/cortical_tiles-2026/crops/2mm/.
If your Morphologist graphs are stored under a non-default sub-path, override it:
pixi run champollion-cortical-tiles \
/data/myproject/derivatives/morphologist-6.0/subjects \
/data/myproject/derivatives/ \
--path_to_graph "t1mri/default_acquisition/*/folds/3.1" \ # glob pattern to .arg graph files
--path_sk_with_hull "t1mri/default_acquisition/default_analysis/segmentation" # skeleton directoryThe --path_to_graph option supports wildcards (*) for variable path segments.
Verify that 28 sulcal region folders were created:
ls /data/myproject/derivatives/cortical_tiles-2026/crops/2mmTo skip subjects with failing quality control, pass a tab-separated file with participant_id and qc columns (1 = keep, 0 = skip):
participant_id qc comments
sub-001 1
sub-002 0 Motion artefact
pixi run champollion-cortical-tiles ... --sk_qc_path /path/to/qc.tsv| Version | Description |
|---|---|
canonical_25 |
Original masks, used for Champollion V1. |
canonical_corrected_26_1 |
Revised labeling with reduced artifacts. |
Pass --masks canonical_25 (or another version) to override the default.
All options
| Option | Description |
|---|---|
--sk_qc_path |
Path to QC TSV file |
--njobs |
Number of CPU cores (default: auto) |
--region-file |
Custom sulcal region configuration file |
--input-types |
Input types to generate (e.g. skeleton foldlabel extremities). Default: all. |
--skip-distbottom |
Skip distbottom generation (saves time; not needed for inference) |
--masks |
Mask version tag |
--regions |
Restrict to specific sulcal regions (space-separated) |
Create the YAML configuration files that link your dataset's crop paths to the Champollion model:
pixi run champollion-config \
/data/myproject/derivatives/cortical_tiles-2026/crops/2mm \ # path to 2mm crops (step 2 output)
--dataset myproject \ # dataset name used in YAML and paths
--output /data/myproject/derivatives/champollion_V1/configs # config root; YAMLs land at {output}/dataset/{dataset}/This writes region YAML files to configs/dataset/myproject/ and sets dataset_folder in dataset_localization/local.yaml so the model knows where your data lives. Pass the same --output path as --config_path in the next step (plus /dataset/myproject).
If your crops live outside derivatives/cortical_tiles-2026/ (e.g. a legacy dataset under deep_folding-2025/), add --external_crops:
pixi run champollion-config \
/data/myproject/derivatives/deep_folding-2025/crops/2mm \ # legacy crop path
--dataset myproject \
--external_crops \ # use the exact crop path instead of the standard derivatives layout
--output /data/myproject/derivatives/champollion_V1/configsWhen the pipeline directory is read-only, write local.yaml to a writable path with --external-config:
pixi run champollion-config \
/path/to/crops/2mm \
--dataset myproject \
--output /writable/path/configs \
--external-config /writable/path/configs/dataset_localization/local.yaml # write local.yaml here instead of inside the pipeline dirAll options
| Option | Description |
|---|---|
--champollion_loc |
Path to Champollion binaries (default: external/champollion_V1) |
--output |
Configs root directory. Region YAMLs land at {output}/dataset/{dataset}/. |
--external_crops |
Use the exact crop path instead of assuming the standard derivatives layout. |
--external-config |
For read-only containers: write local.yaml to a writable path. |
Run the pre-trained Champollion encoders across all 56 model folds (28 regions × 2 hemispheres). Models are downloaded automatically from Hugging Face on first run and cached locally.
pixi run champollion-embeddings \
neurospin/Champollion_V1 \ # model source: HF repo ID, local path, or archive
local \ # localization preset (matches local.yaml written in step 3)
myproject \ # dataset name (must match --dataset from step 3)
my_run \ # run label used to name output CSV files
--embeddings_only \ # skip classifier training, compute embeddings only
--config_path /data/myproject/derivatives/champollion_V1/configs/dataset/myproject # step 3 outputEach of the 56 folds writes a full_embeddings.csv (one row per subject, columns = embedding dimensions) under:
models_cache/Champollion_V1/{region}/
my_run_random_embeddings/full_embeddings.csv
To re-run on an existing dataset, add --overwrite.
You can point to different model sources:
| Source | Example value |
|---|---|
| Hugging Face repo | neurospin/Champollion_V1 |
| Cached local directory | /data/myproject/derivatives/champollion_V1/models_cache/Champollion_V1 |
| Local archive | /path/to/models.tar.gz |
| Remote archive URL | https://example.com/models.tar.gz |
When using Hugging Face, models are cached in data/{dataset}/derivatives/champollion_V1/models_cache/. Pass the cached path directly on subsequent runs to avoid network checks. Use --no-cache to force a full re-download.
All options
| Option | Description |
|---|---|
--config_path |
Path to dataset config directory (generated in step 4) |
--embeddings_only |
Only compute embeddings (skip classifier training) |
--cpu |
Force CPU usage (disable CUDA) |
--overwrite |
Overwrite existing embeddings |
--no-cache |
Force re-extraction of archive |
--run-cka |
Run CKA coherence test after embeddings |
--split |
Splitting strategy: random or custom (default: random) |
--nb_jobs |
Number of CPU workers for the DataLoader |
--cortical_version |
Override the cortical tiles folder name (e.g. cortical_tiles-2025) |
--legacy |
Rewrite config YAML paths to use deep_folding-2025 (for older datasets) |
--labels |
Labels for classifiers (default: ['Sex']) |
--classifier_name |
Classifier type (default: svm) |
Collect the 56 per-region embedding CSVs into a single output directory:
pixi run champollion-combine \
--path_models /data/myproject/derivatives/champollion_V1/models_cache/Champollion_V1/ \ # models cache dir (step 4 output)
--embeddings_subpath my_run_random_embeddings/full_embeddings.csv \ # {run_label}_{split}_embeddings/full_embeddings.csv
--output_path /data/myproject/derivatives/champollion_V1/embeddings/ # destination for the 56 combined CSVs--embeddings_subpath is {short_name}_{split}_embeddings/full_embeddings.csv, using the my_run label and random split from step 5.
Verify 56 CSV files were created:
ls /data/myproject/derivatives/champollion_V1/embeddings/*.csv | wc -lGenerate visualizations: sulcal graph meshes, cortical tile masks, and UMAP scatter plots projecting your subjects onto a pre-trained reference embedding space.
pixi run champollion-snapshots \
--embeddings_dir /data/myproject/derivatives/champollion_V1/embeddings/ \ # step 5 output (56 CSVs)
--reference_data_dir reference_data/ \ # pre-trained UMAP models (in repo)
--output_dir /data/myproject/derivatives/champollion_V1/snapshots/ # where to write imagesAdd --morphologist_dir and --cortical_tiles_dir to also generate mesh and mask snapshots:
pixi run champollion-snapshots \
--morphologist_dir /data/myproject/derivatives/morphologist-6.0/ \ # for sulcal graph mesh snapshots
--cortical_tiles_dir /data/myproject/derivatives/cortical_tiles-2026/crops/2mm/ \ # for tile mask snapshots
--embeddings_dir /data/myproject/derivatives/champollion_V1/embeddings/ \
--reference_data_dir reference_data/ \
--output_dir /data/myproject/derivatives/champollion_V1/snapshots/Use --sulcal-only, --tiles-only, or --umap-only to generate only one snapshot type.
UMAP scatter plots project each subject's sulcal embeddings onto pre-trained 2D maps (one per region and hemisphere). Each plot shows a reference cloud of embeddings from a large pre-trained cohort, with the new subject highlighted. Pre-trained UMAP artifacts are stored in reference_data/ and contain no subject identifiers.
By default all regions with an available embedding CSV and a pre-trained model are plotted. Restrict to specific regions with --umap_region:
pixi run champollion-snapshots \
--embeddings_dir /path/to/embeddings/ \
--reference_data_dir reference_data/ \
--output_dir /path/to/snapshots/ \
--umap-only --umap_region FColl-SRh # restrict to a single regionIf a subject has several Morphologist acquisitions (e.g. two time points), the script warns and uses the first one found. Specify the acquisition explicitly to avoid ambiguity:
pixi run champollion-snapshots \
--morphologist_dir /path/to/subjects/ \
--subject sub-001 --acquisition wk40 \ # acquisition label to disambiguate multiple time points
--output_dir /path/to/snapshots/ --sulcal-onlyAll options
| Option | Description |
|---|---|
--morphologist_dir |
Path to Morphologist output (for sulcal graph snapshots) |
--subject |
Subject folder name (default: first subject found) |
--acquisition |
Acquisition label when a subject has multiple segmentations |
--cortical_tiles_dir |
Path to crops/2mm/ (for tiles mask snapshots) |
--embeddings_dir |
Path to combined embeddings (for UMAP scatter plots) |
--reference_data_dir |
Path to pre-trained UMAP models and reference coordinates |
--umap_region |
Comma-separated region name(s) to plot |
--output_dir |
Directory to save snapshot images |
--sulcal-only / --tiles-only / --umap-only |
Generate only one snapshot type |
--width / --height |
Snapshot dimensions (default: 800×600) |
--tiles_level |
Cortical tiles level to visualize |
champollion_pipeline/
├── external/
│ ├── champollion_V1/ # Self-supervised encoder submodule (contrastive learning)
│ └── cortical_tiles/ # Sulcal region crop extraction submodule
├── reference_data/ # Pre-trained UMAP models and anonymous reference coordinates
├── src/
│ └── champollion_pipeline/ # Installable Python package (entry points: champollion-*)
│ ├── generate_morphologist_graphs.py
│ ├── run_cortical_tiles.py
│ ├── generate_champollion_config.py
│ ├── generate_embeddings.py
│ ├── put_together_embeddings.py
│ ├── generate_snapshots.py
│ └── train_champollion.py
├── data/ # Created by install-all; not committed
└── pixi.toml
pixi run test # All tests
pixi run test-unit # Unit tests only
pixi run test-integration # Integration tests only
pixi run test-smoke # Smoke tests only
pixi run test-cov # Tests with coverage report
pixi run test-fast # Stop on first failureAll dependencies are managed through pixi.toml. Core requirements:
- Python ≥ 3.8
- PyTorch
- BrainVISA / Morphologist (for step 2 only — see brainvisa.info)
- huggingface-hub