Metadata-Version: 2.5
Name: publicdata-au
Version: 0.1.0
Summary: Query and download Australian government open data from publicdata.au.
Project-URL: Homepage, https://publicdata.au/
Project-URL: Documentation, https://publicdata.au/agents/
Project-URL: Source, https://github.com/National-Digital/publicdata.au/tree/main/clients/python
Author: National Digital
License-Expression: MIT
License-File: LICENSE
Keywords: australia,government data,open data,publicdata.au
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Provides-Extra: pandas
Requires-Dist: pandas>=2.0; extra == 'pandas'
Requires-Dist: pyarrow>=14; extra == 'pandas'
Description-Content-Type: text/markdown

# publicdata-au

Query and download Australian government open data from [publicdata.au](https://publicdata.au/).
publicdata.au republishes datasets that governments already publish under open licences, keeps
every version at a URL that never changes, and serves each one in ten formats with a query API.

This package works for every dataset the site serves, named by its slug, so a dataset added to
the site needs no new release.

```
pip install publicdata-au            # queries and downloads, no dependencies
pip install "publicdata-au[pandas]"  # adds read() into a pandas DataFrame
```

## Find a dataset

```python
import publicdata_au as pd_au

pd_au.datasets("road crashes")  # slug, title, publisher, licence and page of each match
pd_au.versions("au-road-deaths")  # every version kept, newest first
```

## Query rows and totals

```python
from publicdata_au import gte, in_

deaths = pd_au.rows(
    "au-road-deaths",
    {"state": in_("QLD", "NSW"), "year": gte(2020)},
    select=["state", "year", "road_user"],
    order="year.desc",
    all=True,
)
deaths.version  # the version the rows came from
deaths.attribution  # the attribution the publisher's licence requires
deaths.to_pandas()

pd_au.aggregate("au-road-deaths", group="state", metric="count", where={"year": 2025})
```

A plain value must match exactly, a list matches any of its values and None matches a blank or
suppressed cell. The filters are `eq`, `neq`, `gt`, `gte`, `lt`, `lte`, `like`, `ilike`, `in_`,
`is_null` and `not_`.

Without `version=` an answer comes from the newest version and changes when the publisher
releases again. Pass a date from `versions()` for an answer that never changes.

## Whole tables

```python
df = pd_au.read("au-road-deaths")  # needs the [pandas] extra
df.attrs["publicdata"]  # the version, licence and attribution the file itself carries
pd_au.download(
    "au-road-deaths", "csv"
)  # parquet, csv, csv.gz, json, ndjson, sqlite, xlsx, arrow, geojson, gpkg
```

Files have no rate limit. The query API allows 60 requests in 10 seconds from one address, and
this package waits and retries when it answers 429.

## Licence and attribution

The data is under each publisher's own licence, which requires the attribution string that every
answer carries. Please also name publicdata.au and link to the version you used. publicdata.au is
an independent republication, and the publishers have not endorsed it.

The package itself is under the MIT licence.
