Pandera

Pandera

Statistical data testing for dataframes

Description

A column quietly fills with nulls or wrong types and downstream models miscompute for weeks. Pandera is a lightweight, flexible statistical data testing library for dataframes.

Declare column types, ranges and statistical checks in schemas that validate as data flows, for pandas, Polars and PySpark.

Features



Schemas:Types and ranges.

Statistics:Hypothesis checks.

Frameworks:pandas, Polars, PySpark.

Early:Catch in pipelines.