Metadata-Version: 2.5
Name: publicdata-au
Version: 0.3.0
Summary: Query and download Australian government open data from publicdata.au.
Project-URL: Homepage, https://publicdata.au/
Project-URL: Documentation, https://publicdata.au/agents/
Project-URL: Source, https://github.com/National-Digital/publicdata.au/tree/main/clients/python
Author: National Digital
License-Expression: MIT
License-File: LICENSE
Keywords: australia,government data,open data,publicdata.au
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Provides-Extra: duckdb
Requires-Dist: duckdb>=1.1; extra == 'duckdb'
Provides-Extra: geo
Requires-Dist: geopandas>=0.14; extra == 'geo'
Requires-Dist: pyogrio>=0.7; extra == 'geo'
Provides-Extra: pandas
Requires-Dist: pandas>=2.0; extra == 'pandas'
Requires-Dist: pyarrow>=14; extra == 'pandas'
Description-Content-Type: text/markdown

# publicdata-au

Query and download Australian government open data from [publicdata.au](https://publicdata.au/).
publicdata.au republishes datasets that governments already publish under open licences, keeps
every version at a URL that never changes, and serves each one in twelve formats with a query API.

This package works for every dataset the site serves, named by its slug, so a dataset added to
the site needs no new release.

```
pip install publicdata-au            # queries and downloads, no dependencies
pip install "publicdata-au[pandas]"  # adds read() into a pandas DataFrame
pip install "publicdata-au[duckdb]"  # adds connect() and relation(), DuckDB on a version
pip install "publicdata-au[geo]"     # adds read_geo() into a geopandas GeoDataFrame
```

## Find a dataset

```python
import publicdata_au as pd_au

pd_au.datasets("road crashes")  # slug, title, publisher, licence and page of each match
pd_au.datasets(topic="roads", jurisdiction="Qld")
pd_au.datasets(publisher="Bureau of Meteorology")
pd_au.versions("au-road-deaths")  # every version kept, newest first
```

## Query rows and totals

```python
from publicdata_au import gte, in_

deaths = pd_au.rows(
    "au-road-deaths",
    {"state": in_("QLD", "NSW"), "year": gte(2020)},
    select=["state", "year", "road_user"],
    order="year.desc",
    all=True,
)
deaths.version  # the version the rows came from
deaths.attribution  # the attribution the publisher's licence requires
deaths.to_pandas()

pd_au.aggregate("au-road-deaths", group="state", metric="count", where={"year": 2025})
```

A plain value must match exactly, a list matches any of its values and None matches a blank or
suppressed cell. The filters are `eq`, `neq`, `gt`, `gte`, `lt`, `lte`, `like`, `ilike`, `in_`,
`is_null` and `not_`.

Without `version=` an answer comes from the newest version and changes when the publisher
releases again. Pass a date from `versions()` for an answer that never changes.

## Whole tables

```python
df = pd_au.read("au-road-deaths")  # needs the [pandas] extra
df.attrs["publicdata"]  # the version, licence and attribution the file itself carries
pd_au.download("au-road-deaths", "csv")  # or parquet, csv.gz, json, ndjson, sqlite, duckdb, ...
```

A version never changes once published, so a downloaded file can be kept and reused. Nothing is
kept unless you ask, with `cache=True` on a call or `PUBLICDATA_CACHE=1` for every call:

```python
df = pd_au.read("au-road-deaths", cache=True)  # downloaded once, read from disk after
pd_au.cache_list()  # what is kept, in cache_dir()
pd_au.cache_clear("au-road-deaths")
```

## Query a version in place

Every version has a DuckDB file, and `connect()` attaches it read-only over HTTPS. Only the
blocks a query touches are read, so a count over millions of rows runs without a download.

```python
con = pd_au.connect("au-road-deaths")  # needs the [duckdb] extra
con.sql("SELECT state, count(*) FROM records GROUP BY 1").df()
con.publicdata  # the version, its URL, licence, attribution and citation

r = pd_au.relation("au-road-deaths")  # one table as a lazy DuckDB relation
r.filter("year >= 2020").aggregate("state, count(*) AS n").df()
```

A database such as G-NAF is one DuckDB file holding every table, the keys between them and the
publisher's views, with one Parquet file per table beside it.

```python
pd_au.tables("gnaf")  # every table with its fields, keys and references
con = pd_au.connect("gnaf")
con.sql("SELECT postcode, count(*) AS n FROM address_view GROUP BY 1 ORDER BY 2 DESC LIMIT 10").df()
pd_au.read("gnaf", table="locality")  # one table as a pandas DataFrame
pd_au.download("gnaf", table="state")  # one table as Parquet
```

For heavy work on a large database, `connect("gnaf", cache=True)` downloads the file once and
queries it from disk. G-NAF's DuckDB file is about 3 GB.

## Maps

Datasets with a location or a shape, such as crash points or ABS boundaries, have a GeoPackage,
which `read_geo()` reads in the reference system the publisher used, usually GDA2020
(EPSG:7844).

```python
lga = pd_au.read_geo("abs-lga-2025")  # needs the [geo] extra
lga.plot()
```

## What changed

Each release is compared with the one before it on the dataset's key.

```python
pd_au.changes("rba-money-market-daily")  # one entry per release: added, removed, changed
pd_au.diff("rba-money-market-daily")  # the newest release in full, with the keys
```

Files have no rate limit. The query API allows 60 requests in 10 seconds from one address, and
this package waits and retries when it answers 429 or a passing server error.

## Licence and attribution

The data is under each publisher's own licence, which requires the attribution string that every
answer carries. Please also name publicdata.au and link to the version you used; `cite(slug)` gives
the citation and `cite(slug, format="bibtex")` the BibTeX entry. Where a licence sets a condition
beyond attribution, such as G-NAF's rule on mail compilation, the package warns once per dataset
with a `LicenceCondition` warning. publicdata.au is
an independent republication, and the publishers have not endorsed it.

The package itself is under the MIT licence.
