Metadata-Version: 2.5
Name: ad-data-pf
Version: 0.2.0
Summary: Tables and figures for a job-ad amenities project, run on Statistics Denmark microdata
Author-email: AskerNC <hms467@econ.ku.dk>
License-Expression: MIT
License-File: LICENSE
Requires-Python: >=3.12
Requires-Dist: loguru>=0.7
Requires-Dist: matplotlib>=3.10
Requires-Dist: numpy>=2.3
Requires-Dist: polars>=1.43
Description-Content-Type: text/markdown

# ad-data-pf

Tables (and later figures) for a research project on job-ad amenities, run on
Statistics Denmark (DST) microdata. The functions are written and tested
locally against simulated data, then `pip install`ed on the DST research
server, where the data are cleaned and the functions called.

All numbers in this package come from simulated data.

## Use

```python
from ad_data_pf import validation as val, save_tex, save_fig, setup_log

setup_log(log_folder, 'descriptives')      # logs the package version first
frames = {'All AKU\nrespondents': aku, 'All\njob ads': jobads,
          'Linked\njob ads': jobads_linked, 'Job ads linked\nwith AKU response': aku_linked}
save_tex(val.desc_table(frames), out / 'tab_base_rates.tex')
save_tex(val.omission_table(aku_linked, ci='wilson'), out / 'tab_validation_measures.tex')
save_fig(val.prevalence_figure(aku_linked, aku), out / 'fig_sensitivity_occ_prevalence.pdf')
```

Tables return a bare `tabular`; figures return a matplotlib `Figure`.
Defaults (variables, labels, panel titles) sit at the top of each module and
in `ad_data_pf.labels`, and can be overridden per call. See the docstrings,
and `examples/` for every exhibit and for custom tables.

## Arguments

Each function fixes the layout of its exhibit: the panels, and the statistic
each one computes. The arguments say only which columns go in and what they
are called, in three plain forms:

- `{column: label}`, one row per column:
  `t_vars={'any_T': 'Any non-standard hours', 'work_night': 'Night work'}`.
- `{column: splits}`, one row per group of a column:
  - `{code: label}` matches the leading characters of a string column
    (`{'1': 'Managers', '2': 'Managers'}` on the 6-digit `disco`) and the
    value of any other (`{True: 'Public', False: 'Private'}`). Codes may
    share a label.
  - `[(upper bound, label), ...]` bins a numeric column, left-closed, `None`
    for no bound (`SIZE`).
- `{label: frame}` where the columns, rows or points are samples
  (`desc_table`, `robustness_table`, `ladder_figure`). The label is the
  column header, row label or axis label, and `\n` breaks it into lines.
  Make each frame with polars before the call:
  `al.filter(c.months_since_hire <= 3)`, `al.with_columns(any_T=c.any_T_sometimes)`.

Panel titles can be renamed and `{}` drops a panel. There are no
package-specific spec objects and no polars expressions as arguments.

`validation` makes these exhibits:

| Exhibit | Function |
| --- | --- |
| Samples and base rates | `desc_table` (panels units, T, M, composition; any number of frames) |
| Confusion table | `confusion_table` (shares of N by default, or counts) |
| Validation measures by type | `omission_table` (Wilson or cluster-bootstrap intervals) |
| Measure by occupational prevalence | `prevalence_figure` (`measure='sens'` or `'fom'`) |
| Measures by group | `heterogeneity_table` (one panel per split column) |
| Measures by linkage criterion | `ladder_figure` |
| Robustness | `robustness_table` (one row per sample) |

Inputs are checked. The following raise an error:

- T or M that is not Boolean, or has nulls;
- a row whose columns exist in no frame;
- a split that no row falls in, or a group with no pairs;
- T and M labels that do not pair up.

Cells resting on fewer than 5 observations are left empty (DST disclosure).
In splits, a null or NaN counts as missing: in polars NaN compares above
every number, so it would otherwise land in the top bin.

On the server, install an exact version (`pip install ad-data-pf==0.2.0`) so a
rerun reproduces the same table. [CHANGELOG.md](CHANGELOG.md) lists what to
retype when moving to a new version.

## Layout

```text
src/ad_data_pf/
├── functions/          one module per topic: validation.py, ...
│                       imported from the top: from ad_data_pf import validation
├── simulation/         fake data per module (validation.py) and the files in data/
└── examples/
    └── validation/     code/ runs every table on the fake data;
                        logs/ and output/ hold what it writes
```

`simulation/` and `examples/` are installed with the package. To find them:

```bash
python -c "import ad_data_pf, pathlib; print(pathlib.Path(ad_data_pf.__file__).parent)"
python <that folder>/examples/validation/code/run_validation.py <output folder>
```

Every example also runs in the VS Code Interactive window (or any Jupyter
kernel), whole or cell by cell (`# %%`). It writes to its own `logs/` and
`output/`, or to `./ad_data_pf_examples/<name>/` in the working directory if
the installed package is read-only.

## Adding a module

1. `functions/<name>.py`: importable at once as `ad_data_pf.<name>`;
   arguments in the forms of [Arguments](#arguments)
2. `simulation/<name>.py`: `simulate()`, `save()`, `load()`
3. `examples/<name>/code/run_<name>.py`, plus `logs/.gitkeep`. Take the
   output folder from `root = example_root('<name>')`, never from `sys.argv`
   or `__file__`, and split sections with `# %%`. `tests/test_examples.py`
   runs every example as a script and as in a Jupyter kernel.
4. `tests/test_<name>.py`

## Development

The package lives in `ad_data_pf/` of the (private) project repository; run
everything below from that folder.

```bash
uv sync                                          # environment in .venv
uv run pytest                                    # also compiles the LaTeX if latexmk is found
uv run python src/ad_data_pf/simulation/validation.py   # save the fake data to simulation/data
```

The lower bounds in `pyproject.toml` are the versions on the DST server. To
test against exactly those, use a separate environment so `uv.lock` is not
rewritten:

```bash
cp uv.lock "$TEMP/uv.lock"
UV_PROJECT_ENVIRONMENT="$TEMP/venv-server" uv sync --resolution lowest-direct
UV_PROJECT_ENVIRONMENT="$TEMP/venv-server" uv run --no-sync pytest
cp "$TEMP/uv.lock" uv.lock
```

GitHub Actions does both on every push that touches `ad_data_pf/`
(`.github/workflows/ad_data_pf_test.yml` at the repository root).

## Release

Bump `version` in `pyproject.toml` and add an entry to
[CHANGELOG.md](CHANGELOG.md). For each change to a call, give the old and
the new call, since the server script is retyped by hand. Commit, then
push a tag `ad_data_pf-v<version>`; GitHub Actions checks that the tag matches the
version, runs the tests, builds and publishes to PyPI (trusted publishing,
`.github/workflows/ad_data_pf_publish.yml`):

```bash
git tag ad_data_pf-v0.1.1 && git push origin ad_data_pf-v0.1.1
```
