Metadata-Version: 2.4
Name: pnadtables
Version: 0.1.0.dev1
Summary: folktables-inspired fairness benchmark tasks on Brazilian PNAD Contínua (IBGE) microdata
Author: Bruno H. M. Oliveira
License-Expression: MIT
Project-URL: Homepage, https://github.com/Brunooliveirab/pnadtables
Project-URL: Issues, https://github.com/Brunooliveirab/pnadtables/issues
Keywords: fairness,machine-learning,benchmark,census,brazil,pnadc,folktables
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas<3,>=1.5
Requires-Dist: numpy>=1.23
Requires-Dist: pnadium==0.25
Provides-Extra: baseline
Requires-Dist: scikit-learn>=1.2; extra == "baseline"
Provides-Extra: examples
Requires-Dist: scikit-learn>=1.2; extra == "examples"
Requires-Dist: shap>=0.42; extra == "examples"
Provides-Extra: dev
Requires-Dist: pytest>=7; extra == "dev"
Requires-Dist: scikit-learn>=1.2; extra == "dev"
Provides-Extra: bigquery
Requires-Dist: basedosdados>=2.0; extra == "bigquery"
Requires-Dist: google-cloud-bigquery>=3.10; extra == "bigquery"
Dynamic: license-file

# pnadtables

[![CI](https://github.com/Brunooliveirab/pnadtables/actions/workflows/ci.yml/badge.svg)](https://github.com/Brunooliveirab/pnadtables/actions)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)

**Fairness benchmark tasks on Brazil's PNAD Contínua (IBGE), with an API inspired by
[folktables](https://github.com/socialfoundations/folktables).** Microdata are accessed
directly from the **official IBGE FTP** through `pnadium`; Base dos Dados / BigQuery is an
optional backend.

*Leia em [português](README.pt-BR.md).*

## Status — pre-release, partially validated against real data

This package is **not ready for use in published results.** Version `0.1.0.dev1` is a
working pre-release: its logic is tested offline, and the SP/2024Q1-Q2 and MT/2024Q1
official files passed real-data validation ([SP Q1](docs/validacao-dados-reais.md),
[SP Q2](docs/validacao-sp-2024t2.md), [MT Q1](docs/validacao-mt-2024t1.md)). The
experimental baseline now covers a household-grouped split
([report](docs/baseline-sp-2024t1.md)), temporal transfer
([report](docs/baseline-temporal-sp-2024t1-t2.md)), and geographic transfer
([report](docs/baseline-geografico-sp-mt-2024t1.md)). The primary baseline also
has conditional 95% intervals from all 200 PNADC bootstrap weights
([report](docs/baseline-sp-2024t1-incerteza.md)). Training uncertainty and
reproduction by a second person remain open.
No release has been tagged and no DOI exists.
For a detailed Portuguese account of the project history, architecture, data and
experiments, see the [complete project report](docs/relatorio-completo-pnadtables.md).

Open gates before `v0.1.0` (see [`pnadtables-plano.md`](pnadtables-plano.md)):

| Gate | What it requires | Status |
|---|---|---|
| G0 | Git repository, correct metadata, no secrets | ✅ |
| G1 | Schema and value domains confirmed against the IBGE dictionary and real data | ✅ SP/2024Q1-Q2 and MT/2024Q1 passed |
| G2 | Deterministic encoding, missingness policy, weight validation, integration tests | 🟨 local validation implemented; real-data integration still open |
| G3 | Leakage-free splits, survey-design protocol, metrics with uncertainty | 🟨 splits and conditional bootstrap intervals done; training uncertainty remains open |
| G4 | Reproducible real-data baseline with a results table | 🟨 real-data reports produced; second-person reproduction remains open |
| G5 | CI green on supported Python versions, clean install from artifact | ⬜ |
| G6 | Datasheet, limitations, provenance, verified licences/terms | ⬜ |
| G7 | Public repo, `v0.1.0` tag, archived release, DOI | ⬜ |
| G8 | PyPI publication, verified `pip install pnadtables` | ⬜ |

```python
from pnadtables import PNADCDataSource, PNADEmployment, audit_report

src = PNADCDataSource(
    ano=2024, trimestre=1, ufs=["MT", "SP"],
    caminho="dados", salvar=True, formato="csv",
    # incluir_pesos_replicados=True adds the 200 bootstrap weights for CIs.
)
X, y, group, weight = PNADEmployment.df_to_pandas(src.get_data())
# 4-tuple: folktables returns (X, y, group); pnadtables adds the PNADC survey weight.
# Code written against folktables will NOT unpack this without modification.
# Nominal features are pandas categoricals; fit an encoder on training data only.
```

`df_to_numpy()` refuses nominal features instead of inventing integer codes. Use
`df_to_pandas()` with a training-only preprocessing pipeline (see `example_audit.py`).

## Why

Most empirical fairness research is calibrated on U.S. Census data (UCI Adult, then
folktables). Whether its conclusions transfer to the Global South is, as far as we know,
largely untested — but **this is a hypothesis we have not yet verified with a systematic
literature search** (ACM DL, IEEE Xplore, Scopus/OpenAlex, arXiv, Google Scholar). Do not
cite this README as evidence that no comparable benchmark exists. `pnadtables` adapts the
`BasicProblem` design to Brazil's official labour survey and targets two things:

- **Racial taxonomies are not translatable 1:1.** IBGE uses five self-declared categories
  (Branca, Preta, Amarela, Parda, Indígena), plus code 9 for "Ignorada" (non-response).
  Collapsing *Preta + Parda* into a binary Black/white axis is a theoretical choice, not a
  data fact. The package ships both encodings so results can be reported under each.
- **Complex survey design.** PNADC is a weighted, stratified, clustered sample. `weight`
  alone yields weighted point estimates. With `incluir_pesos_replicados=True`, the
  baseline uses all 200 bootstrap weights for conditional intervals under the sampling
  design. These intervals keep predictions fixed and do not cover training uncertainty.

### Not a drop-in folktables replacement

Name parity is not construct parity. ACS and PNADC differ in universe, periodicity,
variables and sample design, and the race/colour categories are not semantically
interchangeable. A Brazil–U.S. comparison needs an explicit harmonisation table
(construct, universe, income window, features, categories, geography, design, metric),
not just running tasks with similar names.

## Tasks

| Task | Conceptual reference | Universe | Target |
|---|---|---|---|
| `PNADEmployment` | ACSEmployment | age ≥ 14 | employed (`vd4002 == 1`) |
| `PNADIncome` | ACSIncome | employed, age ≥ 14, income > 0 | monthly income > threshold (parameter) |
| `PNADIncomeRegression` | continuous complement | employed, age ≥ 14, income > 0 | monthly income in BRL |

Features: age, sex, race/colour, education level, state (UF). Protected attribute:
race/colour. Build new tasks with `BasicProblem(features, target, target_transform, group, ...)`.

## Install

```bash
pip install git+https://github.com/Brunooliveirab/pnadtables
pip install "pnadtables[baseline]"   # adds scikit-learn for pnadtables-baseline
pip install "pnadtables[examples]"   # adds scikit-learn + shap for example_audit.py
```

The default backend downloads the public quarterly file from IBGE through `pnadium`. It
requires no account, credentials or billing project. Selecting fewer variables reduces
memory use, although the compressed quarterly archive still has to be downloaded.

```python
from pnadtables import PNADCDataSource

src = PNADCDataSource(ano=2024, trimestre=1, ufs=None)  # all 27 states
df = src.get_data()
```

Run the reproducible baseline with:

```bash
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet

# Continuous income: log1p linear fit, evaluated in BRL with MAE/RMSE.
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet \
  --tarefa income_regression

# Train on SP/2024Q1 and evaluate on SP/2024Q2, removing repeated households.
pnadtables-baseline dados/pnadtables_pnadc_2024T1_SP.parquet \
  --teste-externo dados/pnadtables_pnadc_2024T2_SP.parquet \
  --saida docs/baseline-temporal-sp-2024t1-t2.md
```

The table below is the historical binary-race result from before `0.1.0.dev1`; reproduce
it with `--race-mode binary`. The current default uses five groups.

| Model | Weighted accuracy (95% CI) | Weighted ROC AUC (95% CI) |
|---|---:|---:|
| Weighted majority | 0.617 [0.604, 0.630] | 0.500 [0.500, 0.500] |
| Logistic regression with protected attribute | 0.690 [0.678, 0.702] | 0.721 [0.708, 0.734] |
| Logistic regression without protected attribute | 0.690 [0.678, 0.701] | 0.716 [0.703, 0.730] |

These SP/2024Q1 intervals use the 200 bootstrap weights and condition on the fitted
predictions; they do not include training uncertainty. See the
[methodology](docs/metodologia-incerteza.md).

### Optional BigQuery backend

Install `pnadtables[bigquery]` and use the explicit backend when remote columnar querying
is preferable. This path requires a billed Google Cloud project. It performs a free dry
run and enforces `maximum_bytes_billed` before executing the query.

```python
from pnadtables import PNADCBigQueryDataSource

src = PNADCBigQueryDataSource("my-gcp-project", 2024, 1, ufs=["MT", "SP"])

print(src.dry_run())     # free: what this query would cost, before running it
df = src.get_data()      # refuses above DEFAULT_MAX_GB (5 GB)
df = src.get_data(max_gb=20)   # raise the ceiling deliberately
df = src.get_data(max_gb=None) # disable both guards — know what you are doing
```

See [BigQuery pricing](https://cloud.google.com/bigquery/pricing) and
[cost controls](https://docs.cloud.google.com/bigquery/docs/best-practices-costs). Never
replace the selected column list with `SELECT *`.

## Before you publish numbers

1. Run `pnadtables-validar-dicionario --snapshot docs/pnadc-dicionario.snapshot.json`
   for the versioned evidence, or use `--escrever` to refresh it from IBGE. Validate a
   downloaded quarter's domains before publishing. If using the optional BigQuery backend,
   also run `inspect_bigquery_schema(billing_project_id)` to verify that ingestion.
   Run `pnadtables-validar-dados <arquivo.parquet>` on every experimental vintage.
2. Set the income threshold in `make_pnad_income(threshold=...)` to the reference year of
   your data.
3. Use the five self-declared race/colour categories as the primary analysis (the API and
   CLI default), then run `--race-mode binary` as a White vs Black (Black + Brown)
   sensitivity analysis and quantify exclusions.
4. Report **both** weighted and unweighted metrics as a sensitivity analysis. Neither is
   universally the correct headline number — say which question each one answers.
5. Do not use a naive random split: people in the same household are correlated and PNADC
   rotates households. Use the household-grouped baseline or `--teste-externo` for a
   temporal/geographic evaluation; repeated `cod_fam` values are removed from the test set.

See [DATASHEET.md](DATASHEET.md) (Gebru et al. structure) for the full methodological record.

## Metrics

`selection_rates`, `disparate_impact`, `rate_by_group` (TPR/FPR), `audit_report` — all with
optional survey weights. Inputs are validated for length, missing/non-binary values and
invalid weights; the report includes raw counts, weight totals, rate differences and
ratios. By default `disparate_impact` uses the highest-rate group as
reference, so ratios fall in (0, 1] and the 0.8 threshold is the relevant one; the
symmetric 1.25 threshold only applies when you pass a fixed `reference` yourself. The 4/5
rule is a contextual screening heuristic, **not** an automatic legal diagnosis of
discrimination.

For the continuous task, `regression_report` reports actual/predicted means, signed error,
MAE and RMSE by group. These are descriptive diagnostics, not causal conclusions. Removing
race/colour from the feature matrix does not remove proxy information carried by geography,
education or other correlated variables.

The baseline is also a direct Python API:

```python
from pnadtables import executar_baseline

result = executar_baseline(df, tarefa="income_regression", race_mode="five")
```

`example_audit.py` runs the full loop: data → model → weighted audit → SHAP attributions.
SHAP is illustrative only; a standalone feature-importance ranking does not support claims
about a "right to explanation".

## Intended use

Research and auditing of algorithmic systems. **Not** for decisions about individuals, and
not a substitute for substantive research on inequality. Do not attempt to re-identify
respondents.

## Development

```bash
pip install -e ".[dev]"
pytest
```
Tests run offline with synthetic PNADC-shaped data and a mocked IBGE download. Passing them
is **not** evidence that every real-data vintage has the expected domains (gate G1).

## Citation

See [CITATION.cff](CITATION.cff). Note that no version has been released yet. Please also
cite folktables:

> Ding, F., Hardt, M., Miller, J., & Schmidt, L. (2021). *Retiring Adult: New Datasets for
> Fair Machine Learning.* NeurIPS 34.

## License

MIT for this package's code. Microdata © IBGE; this package redistributes no data. The
optional backend accesses the Base dos Dados copy. Attribution and applicable access terms
must be verified before publication (gate G6).
