PySUS is a Python package for accessing and analyzing Brazil's public health data (DATASUS). It provides tools to download, process, and work with health datasets including SINAN (disease notifications), SIM (mortality), SINASC (births), SIH (hospitalizations), SIA (ambulatory), CIHA, CNES, PNI, and more.
- Simplified API: New high-level functions for direct DataFrame access
- Streamlit Web UI: Launch a local web interface for browsing and downloading datasets
- Flexible Schema Modes: Read multiple parquet files with union, intersection, or strict modes
- SQL Query: Filter catalog queries by dataset, group, state, year, and month
pip install pysusFor the local Streamlit web interface:
pip install pysus[web]A pre-built JupyterLab image is available on Docker Hub:
docker pull alertadengue/pysus
docker run -p 8888:8888 alertadengue/pysusOr build locally and start the container:
docker compose up --buildThen open http://127.0.0.1:8888/lab in your browser.
Stop the container:
docker compose downBy default, the high-level convenience functions query and download data locally, returning a list of paths to the downloaded Parquet files. This allows you to inspect the file structure or load them with your preferred tool (e.g., pandas, Polars, DuckDB).
frompysusimportsinan, sinasc, sim, sih, sia, pni, ibge, cnes, ciha# Download SINAN Dengue data for 2000 and return a list of Parquet pathsparquet_files=sinan(disease="deng", year=2000)
# Multiple yearsparquet_files=sinan(disease="deng", year=[2023, 2024])
# SINASC births for São Paulo, 2020-2023parquet_files=sinasc(state="SP", year=[2020, 2021, 2022, 2023])
# SIM mortality dataparquet_files=sim(state="SP", year=2024)
# SIH hospitalizations with monthparquet_files=sih(state="SP", year=2024, month=[1, 2, 3])
# CNES health facilitiesparquet_files=cnes(state="SP", year=2024, month=1)If you prefer to load and combine the data automatically into a single pandas DataFrame, pass the as_dataframe=True parameter to any of the functions:
importpandasaspdfrompysusimportsinan# Download and return a concatenated pandas DataFramedf=sinan(disease="deng", year=2024, as_dataframe=True)You can also list the files within the dataset to check which files are available to download
frompysusimportlist_fileslist_files("SINAN")frompysusimportPySUSasyncdefmain():
asyncwithPySUS() aspysus:
# Query DuckLake catalogfiles=awaitpysus.query(
dataset="sinan",
group="DENG",
state="SP",
year=2024,
)
# Download filesforfinfiles:
local=awaitpysus.download(f)
print(local.path)
# Read multiple parquet filesimportglobpaths=glob.glob("/cache/sinan/**/*.parquet")
df=pysus.read_parquet(paths, mode="union")Launch the local web interface:
pysus webOr with a custom port:
pysus web -p 8080Or run directly with Streamlit:
streamlit run pysus/web/app.pyThe web interface provides three data sources:
- Default (DuckLake): Queries the PySUS S3 catalog — the primary data source. Select a dataset and filter by group, state, year, and month.
- FTP DataSUS: Browses legacy DATASUS FTP directories. Auto-connects on tab selection.
- API DataSUS (DadosGov): Queries the dados.gov.br open data API. Requires an API token.
Use the interactive filters to find files, add them to the download queue, and download with a single click. After a query, an expandable Python snippet shows the equivalent code to reproduce the same operation in a script or notebook.
- Automatic Downloads: Fetch data from FTP, DuckLake (S3), and dados.gov.br API
- Parquet Output: All downloaded data is converted to Apache Parquet format
- DuckLake Integration: S3-compatible cloud storage for parquet catalogs
- Local Catalog: SQLite-based tracking of download history to avoid re-downloads
- Type Inference: Automatic data type conversion from legacy formats (DBF, DBC)
- CLI with Streamlit UI: Command-line interface with local web-based UI
PySUS 2.0 has a modular architecture:
PySUS
├── FTP Client # Traditional FTP-based datasets
├── DadosGov Client # dados.gov.br API access
├── DuckLake Client # S3 object storage for Parquet catalogs
└── Database Functions # High-level functions (sinan, sinasc, sim, etc.)
New in PySUS 2.0, these functions provide a simplified interface:
| Function | Dataset | Parameters |
|---|---|---|
sinan(disease, year) | Disease Notifications | disease (e.g., "DENG", "ZIKA"), year |
sinasc(state, year, group) | Births | state, year, group (optional) |
sim(state, year, group) | Mortality | state, year, group (optional) |
sih(state, year, month, group) | Hospitalizations | state, year, month, group (optional) |
sia(state, year, month, group) | Ambulatory | state, year, month, group (optional) |
pni(state, year, group) | Immunizations | state, year, group (optional) |
ibge(year, group) | IBGE | year, group (optional) |
cnes(state, year, month, group) | Health Facilities | state, year, month, group (optional) |
ciha(state, year, month) | Hospital Admissions | state, year, month |
asyncwithPySUS() aspysus:
# Filter by any combination of parametersfiles=awaitpysus.query(
dataset="sinan", # dataset namegroup="DENG", # disease groupstate="SP", # state codeyear=2024, # yearmonth=1, # month (optional)
)# Union mode (default) - includes all columns from any filedf=pysus.read_parquet(paths, mode="union")
# Intersection mode - only common columns across all filesdf=pysus.read_parquet(paths, mode="intersection")
# Strict mode - raises error if schemas don't matchdf=pysus.read_parquet(paths, mode="strict")
# With custom SQLdf=pysus.read_parquet(paths, sql="SELECT * WHERE column > 100")frompysusimportCACHEPATHimportosos.environ['PYSUS_CACHEPATH'] ='/my/custom/path'# orpysus=PySUS(db_path='/my/config.db')PYSUS_CACHEPATH: Directory for cached filesDADOSGOV_TOKEN: API token for the dados.gov.br client (required for DadosGov downloads)ACCESS_KEY/SECRET_KEY: S3 credentials for writing to the PySUS bucket (catalog sync/maintenance)
| Dataset | Description | Source |
|---|---|---|
| SINAN | Disease Notifications | FTP / DuckLake / DadosGov |
| SIM | Mortality | FTP / DuckLake / DadosGov |
| SINASC | Births | FTP / DuckLake / DadosGov |
| SIH | Hospitalizations | FTP / DuckLake |
| SIA | Ambulatory | FTP / DuckLake |
| CIHA | Hospital Admissions | FTP / DuckLake |
| CNES | Health Facilities | FTP / DuckLake / DadosGov |
| PNI | Immunizations | FTP / DuckLake / DadosGov |
| IBGE | Geographic Data | FTP / DuckLake |
conda env create -f conda/dev.yaml
conda activate pysuspoetry installRun code linters:
pre-commit run --all-filesRun tests:
pytest pysus/tests/Run tests inside the Docker container:
docker compose exec -T -w /usr/src jupyter python3 -m pytest pysus/tests/GPL