publicdataau publicdata.au on GitHub
This is an independent republication of open government data. No government agency is affiliated with it or has endorsed it. Every file names its publisher, its licence and the attribution the licence asks for.An independent republication of open government data. No government agency has endorsed it. Read more

Contribute

The code that builds publicdata.au is open source. It is on GitHub under the GNU Affero General Public License, and anyone can propose a change to it. A maintainer at National Digital reviews every pull request, and a merged change goes live with the next release.

National-Digital/publicdata.au

Add a dataset

Any dataset in the backlog with an open licence can be added by anyone. A dataset is one YAML file in the register/ folder, which names the source, the licence with the publisher's own statement as evidence, the attribution and the fields to publish. python -m publicdata register draft writes a first draft from the dataset's portal page. You finish it by hand and build it locally to check it, and the contributing guide has the steps.

The licence is the only thing that stops a dataset. A non-commercial or no-derivatives licence, or none at all, means it cannot be published here, however useful it is.

Add a file format

Each format is a writer that takes one normalised table and its provenance and returns the bytes of the file. A writer reads no network, clock or random value, so two builds of one version give identical files. A new format is added to every dataset at the next deploy. The guide's section on serialisation lists what a writer needs.

Read a new kind of portal

A source adapter fetches from one kind of portal or file host. It reads the licence on every run and dates each version by the publisher's own change date. Datasets on a portal the fetch cannot read yet wait in the backlog, so one adapter can open up many of them. The guide explains how to add one.

Fix something

Every dataset page links to the register entry it is built from, so a clearer description or a field the publisher renamed is a small pull request. Faults in the site, the query API, the MCP server or the Python and R clients go in the issues, along with any file that differs from the publisher's own. Improvements to the documentation are as welcome as code.

Without writing code

Votes in the backlog decide which datasets are built next. A report of a file that differs from the publisher's, or of a page that reads badly, helps as much as a pull request. A publisher that confirms a licence in writing can move a blocked dataset onto the list.

How a change is accepted

Sign off each commit with git commit -s, which certifies under the Developer Certificate of Origin that you may submit it. Title the pull request as a Conventional Commit, such as data(register): add <what it is>. The checks build and test the site and need no credentials, so they run on a pull request from a fork. A maintainer then reviews it and squash-merges it, and the release notes on GitHub name the people whose changes each release carries.

The rules every change is held to are in the guide. The site publishes what the publisher published and derives nothing from it, and a version never changes once it is out. By taking part you agree to the code of conduct. Report a security issue privately, as the security policy describes.

Licences

Each dataset stays under its publisher's licence. The code is under the AGPL and the register's own text is under CC BY 4.0. The name publicdata.au and its mark are outside both licences, and BRAND.md says what a copy of the site may use.