This guide walks you through installing dataprof, running your first profile, and understanding the results. By the end you'll know how to move from a mystery dataset to a report that shows where the data is thin, duplicated, inconsistent, or stale.
Data profiling is the process of examining a dataset to collect statistics and assess quality. It answers questions like:
- How many rows and columns does the data have?
- What data types are present? Are they consistent?
- How many values are missing? Which columns are most affected?
- Are there duplicates? Outliers? Impossible values?
- Is the data fresh or stale?
dataprof automates this analysis and assesses quality across dimensions informed by ISO 8000 and ISO/IEC 25012, producing a structured report you can use for data validation, ETL pipelines, and quality monitoring. The aggregate score is dataprof's configurable formula, not an ISO certification.
Choose the interface that fits your workflow:
# Python
uv pip install dataprof
# or: pip install dataprof
# Rust library -- add to Cargo.toml
# dataprof = "0.10"The published Python wheels cover the base package API for local files, DataFrames, and Arrow objects. If you need async URL profiling or database helpers, build the extension from source with the corresponding Rust features.
import dataprof as dp
report = dp.profile("sales.csv")
# High-level summary
print(f"Rows: {report.rows}")
print(f"Columns: {report.columns}")
print(f"Quality: {report.quality_score}")
# Access columns directly (dict-like)
col = report["price"]
print(f" price: mean={col.mean}, std={col.std_dev}")
# Iterate all columns
for name in report:
col = report[name]
print(f" {col.name}: {col.data_type}, {col.null_percentage:.1f}% null")
# Quick pandas-like summary
print(report.describe())
# Export to pandas, polars, or Arrow for further analysis
df = report.to_dataframe() # pandas
pl_df = report.to_polars() # polars (no pandas needed)
table = report.to_arrow() # pyarrow (no pandas needed)use dataprof::Profiler;
fn main() -> Result<(), dataprof::DataProfilerError> {
let report = Profiler::new().analyze_file("sales.csv")?;
println!("Rows: {}", report.execution.rows_processed);
println!("Columns: {}", report.execution.columns_detected);
if let Some(quality) = &report.quality {
println!("Quality score: {:.1}%", quality.score());
}
for col in &report.column_profiles {
println!(" {} ({:?}): {} nulls out of {}",
col.name, col.data_type, col.null_count, col.total_count);
}
Ok(())
}dataprof evaluates seven quality dimensions informed by ISO 8000 and ISO/IEC 25012. Each is scored from 0 to 100, and dataprof computes its overall score as a configurable weighted average of the dimensions that were actually assessed.
Measures how much data is present vs. missing.
missing_values_ratio-- fraction of null/empty values across all columnscomplete_records_ratio-- fraction of rows with zero nullsnull_columns-- columns past the configured null threshold (50% by default)
The completeness score is on the same 0–100 scale as the other dimensions. It
combines cell completeness (100 - missing_values_ratio) and row completeness
(complete_records_ratio), so inspect both underlying values when setting a
quality gate. complete_records_ratio is deliberately strict: every column is
treated as required, so one sparse optional field can drive it to zero. Inspect
missing_values_ratio and null_columns beside it; a richer optional-column
policy is tracked in #436.
Measures whether values conform to their expected types and formats.
data_type_consistency-- fraction of values matching the column's typeformat_violations-- values that don't match detected patterns (e.g. an email column with non-email values)encoding_issues-- invalid character encoding detected
For a column with an inferred type, data_type_consistency is the fraction of
non-null values conforming to that type. A string column is scored differently,
because every value conforms to string and the metric would report 100% for
any mixture: each value is assigned to one lexical class -- numeric, date,
boolean, or text -- and the score is the share held by the largest class. So a
column of 60% numbers and 40% junk reports 60, while a genuinely textual column
reports 100. Two exceptions keep the plain type check: columns you declared via
identifier_columns, whose schemes mix forms on purpose, and columns whose name
announces dates, which stay held to dates.
data_type_consistency is a dataset-level number and cannot say which column
is mixed. Each column carries its own evidence in type_homogeneity, a count of
non-null values per lexical class:
report["amount_eur"].type_homogeneity
# {"numeric": 6, "date": 0, "boolean": 0, "text": 4}None there means the classification did not run; four zero counts mean it ran
and the column had nothing to classify. The counts cover the values the profiler
retained, so compare their sum against total_count - null_count to see whether
they describe the whole column or the engine's reservoir sample.
Measures duplication in the data.
duplicate_rows-- number of exact duplicate rowskey_uniqueness-- ratio of distinct values in likely-key columnshigh_cardinality_warning-- flagged when a column has suspiciously many unique values
Measures whether values fall within expected ranges.
outlier_ratio-- fraction of values outside the interquartile rangerange_violations-- values outside expected boundsnegative_values_in_positive-- negative numbers in explicitpositive_columns
Measures whether confidently inferred date/time values are current and
consistent. Inferred date columns participate automatically. Pass
temporal_columns=["created_at", "updated_at"] in Python or call
.temporal_columns(...) on the Rust profiler to add columns whose dates cannot
be inferred confidently, such as mixed-format strings.
future_dates_count-- dates that are in the futurestale_data_ratio-- fraction of temporal data that appears outdatedtemporal_violations-- ordering inconsistencies in time series
Measures whether values conform to a confidently detected semantic pattern.
valid_values_ratio-- share matching the dominant patterninvalid_values-- non-null values that do not match itvalues_checked-- values in columns with enough pattern evidence to assess
Columns without a confident pattern are unassessed; dataprof does not invent a domain rule from a weak match.
Measures consistency of effective decimal scale within floating-point columns.
decimal_places_consistency-- share using each column's modal decimal scaleinconsistent_precision_values-- values whose effective scale differsnumeric_values_checked-- parseable finite float values examined
This describes representation consistency. It does not infer how many decimal places a business domain requires.
dataprof auto-detects the delimiter (,, ;, |, \t) by analyzing the first few lines:
report = dp.profile("european_data.csv") # semicolons auto-detected
report = dp.profile("pipe_export.csv") # pipes auto-detectedFor files with inconsistent column counts, enable flexible parsing:
report = dp.profile("messy.csv", csv_flexible=True)These are two different grammars, and dataprof holds you to whichever one you asked for.
.json — a standard JSON document. Exactly one array of objects
([{...}, {...}]) or one object ({...}) as a single record. Whitespace is
insignificant, so the document may be pretty-printed across as many lines as
you like.
.jsonl — JSON Lines. One record per physical line. A record may not span
lines, and a line may not hold more than one value. Blank lines are separators
and are neither records nor errors.
report = dp.profile("users.json")
report = dp.profile("events.jsonl")
# Bytes need the grammar named, since there is no extension to read it from.
report = dp.profile(payload, format="jsonl")The distinction matters most where the two would otherwise disagree. Objects
written back to back with the newline lost — {"a":1}{"b":2} — are not two
records under either grammar: as JSONL the line holds two values, and as JSON
the input holds two documents. Both are reported rather than profiled, because
a concatenation that silently reads as clean data is the failure a profiler
exists to catch.
Every transport applies the same rule: file paths, byte buffers, and the async entry points all agree on what a record is.
Parquet is the most efficient format to profile. Schema inference and row counting read only the file metadata -- zero row scanning required.
schema = dp.infer_schema("data.parquet") # instant schema from metadata
rows = dp.quick_row_count("data.parquet") # instant row count from metadata
report = dp.profile("data.parquet") # full profiling reads row groupsFor very large files, you don't always need to read every row. dataprof provides two mechanisms to limit work.
Profile a representative subset instead of every row:
# Python: a uniform random sample of 100,000 rows
from dataprof import SamplingStrategy
report = dp.profile("huge.csv", sampling=SamplingStrategy.reservoir(100000))
print(report.rows, report.sampling_applied, report.sampling_ratio)// Rust
use dataprof::{Profiler, SamplingStrategy};
let report = Profiler::new()
.sampling(SamplingStrategy::Reservoir { size: 100_000 })
.analyze_file("huge.csv")?;Sampling bounds the cost of analysis, not of reading. A uniform sample can
only be drawn once the source has been seen to the end, so the file is still
read in full and reservoir(n) holds n rows in memory. To read less, use a
stop condition — the two compose.
Sampling applies to CSV sources on the auto and incremental engines and to
every dataprof.asyncio entry point. The columnar engine and the JSON and
Parquet readers cannot sample row by row, so they raise rather than quietly
returning a full profile.
Stop processing when a criterion is met, without needing to know the total size upfront:
from dataprof import StopCondition
# Stop after 10k rows OR 50 MB, whichever comes first
stop = StopCondition.max_rows(10000) | StopCondition.max_bytes(50_000_000)
report = dp.profile("stream.csv", stop_condition=stop)use dataprof::{Profiler, StopCondition};
let report = Profiler::new()
.stop_when(StopCondition::MaxRows(10_000))
.analyze_file("huge.csv")?;When embedding dataprof in a web service or processing live streams, use the async API:
Build the Python extension from source with python-async,async-streaming before using these helpers.
from dataprof.asyncio import profile_file, profile_url
# Profile a local file asynchronously
report = await profile_file("data.csv")
# Profile a remote file over HTTP
report = await profile_url("https://example.com/data.parquet")use bytes::Bytes;
use dataprof::{Profiler, AsyncDataSource, BytesSource, AsyncSourceInfo, FileFormat};
// Profile an async byte stream
let bytes = Bytes::from("name,score\nada,100\n");
let info = AsyncSourceInfo::new("upload", FileFormat::Csv)
.size_hint(Some(bytes.len() as u64));
let source = BytesSource::new(bytes, info);
let report = Profiler::new().profile_stream(source).await?;Profile data directly from PostgreSQL, MySQL, or SQLite without exporting to files.
Build the Python extension from source with python-async,database and the connector feature you need before using this API.
report = await dp.analyze_database_async(
"postgres://user:pass@localhost/mydb",
"SELECT * FROM users",
calculate_quality=True,
)use dataprof::Profiler;
async fn profile_users() -> Result<(), dataprof::DataProfilerError> {
let report = Profiler::new()
.connection_string("postgres://user:pass@localhost/mydb")
.analyze_query("SELECT * FROM users")
.await?;
println!("Profiled {} rows", report.execution.rows_processed);
Ok(())
}See the Database Connectors Guide for connection strings, SSL, retry configuration, and more.
- Examples Cookbook -- copy-pasteable recipes for common tasks
- Python API Guide -- full API reference
- Database Connectors -- advanced database setup
- Contributing -- how to contribute to dataprof