NOTICE-DATA -- licensing framing for data/, docs/, INCIDENTS.md, and
mappings/

Copyright (c) 2026 Emmanuel G. Jr. and contributors.

This file explains how the CC BY 4.0 grant in LICENSE-DATA applies to
this repository's dataset and other non-code material, and records the
specific exceptions to it. It is explanatory only: it grants no rights
beyond, and takes nothing away from, LICENSE-DATA and
docs/SOURCE_LICENSES.md, which remain the operative texts.

This repository is dual-licensed. The dataset and other non-code
material (files under data/, the INCIDENTS.md rollup, and the
mappings/ reference tables) is Creative Commons Attribution 4.0
International (CC BY 4.0), for original and aggregated genai_incidents
content, per LICENSE-DATA. The code in this repository (notably
scripts/ and schema/) is licensed separately under the MIT License --
see LICENSE.

Content sourced from upstream providers carries its own terms in
addition to this grant. The complete, per-source enumeration of
upstream license, scrape-permission, redistribution, and
relicense-compatibility status is recorded in docs/SOURCE_LICENSES.md
and governs wherever it is more restrictive than the CC BY 4.0 grant in
LICENSE-DATA.

As of this writing, three upstream sources carry a content obligation
detailed in full below: the AIAAIC Repository, the AI Incident
Database (AIID), and the OECD AI Incidents and Hazards Monitor (AIM).
This list is maintained by hand and is added to as new sources
resolve; if a source you are looking for is not named here,
docs/SOURCE_LICENSES.md's own summary-of-outcomes table is the
authoritative, always-current enumeration and takes precedence over
this file falling behind it.

Two verbatim text bodies embedded in this dataset are Apache License
2.0, not the CC BY 4.0 grant in LICENSE-DATA, and are not relicensed by
it:

  - MITRE ATLAS case-study text, reproduced verbatim by
    scripts/ingest_external.py, is Copyright 2021-2026 The MITRE
    Corporation, licensed under Apache License 2.0. See
    https://www.apache.org/licenses/LICENSE-2.0 and the ATLAS row of
    docs/SOURCE_LICENSES.md.

  - garak probe-docstring text, concatenated verbatim by
    scripts/ingest_external.py, is Copyright 2023 Leon Derczynski and
    NVIDIA, licensed under Apache License 2.0. See
    https://www.apache.org/licenses/LICENSE-2.0 and the garak row of
    docs/SOURCE_LICENSES.md.

The AIAAIC Repository is a CC BY-SA 4.0 source (share-alike). Per
project decision D2, AIAAIC's own narrative/descriptive cell text is
never carried in this dataset -- what is kept is the AIAAIC-authored
title (a short headline, not narrative) plus short taxonomy-tag facts
(system, technology, sector, jurisdiction, affected) plus a source
pointer to the specific AIAAIC entry -- which disposes of the
*copyright* share-alike question (the separate database-right
question is addressed below; see E13) and keeps this repository's
own LICENSE-DATA grant a single, clean CC BY 4.0 with no
dataset-wide BY-SA carve-out. The retained headline is itself
AIAAIC-authored text, not a bare fact; it is covered, on every entry
that cites an AIAAIC source -- both the rows whose description derives
from AIAAIC's sheet (recorded via description_source == "aiaaic") and
the smaller hand-curated set whose title and categorical facts derive
from AIAAIC directly, not via that field -- by the same row-level CC
BY-SA 4.0 attribution/share-alike marker described below. The current
count of AIAAIC-citing rows is audited, not restated here to avoid
drift, in docs/SOURCE_LICENSES.md section 1.1. (A related, open
question -- whether the headline's own wording, being an editorial
characterisation rather than a bare fact, defeats the "disposes of
the copyright share-alike question"
conclusion above under UK/EU law -- is tracked as escalation E15/D17;
see docs/SOURCE_LICENSES.md section 1.1 and
docs/audits/WS0-E13-database-right-2026-07-18.md section 5.1 item 8.)

That does not fully resolve AIAAIC's licensing status. A separate
EU/UK *sui generis* database-right question over the scale of this
project's AIAAIC extraction remains OPEN, pending AIAAIC's reply to
outreach sent 2026-07-27 (or a qualified-counsel resolution) -- see
docs/audits/WS0-E13-database-right-2026-07-18.md (E13). Per decision
D11, genai_incidents itself holds no UK/EU sui generis database right
(its sole maker is an individual habitually resident in Canada, so the
regulation-18/Article-11 maker-nexus test fails) -- so even if
AIAAIC's own database right subsists and this project's extraction is
judged substantial, the worst case is row-level CC BY-SA 4.0
attribution/share-alike on the specific AIAAIC-derived rows, never a
dataset-wide obligation on LICENSE-DATA as a whole. That row-level
worst case is honored proactively, not merely accepted: every entry
that cites an AIAAIC source -- the rows whose description derives from
AIAAIC's sheet (recorded via description_source == "aiaaic") plus the
hand-curated rows whose title and categorical facts derive from
AIAAIC directly -- carries a machine-readable content_license marker -- source, license, attribution,
obligations -- naming AIAAIC Repository as the attribution target and
share-alike as an honored obligation on that row. The marker travels with the row
wherever it is published: data/incidents.json,
data/incidents.min.json, the HuggingFace export, and the STIX bundle
(as x_content_license). See docs/SOURCE_LICENSES.md section 1.1 for
full detail.

The AI Incident Database (AIID) is a CC BY-SA 4.0 source (share-alike)
for its incidents/quickadd/duplicates/taxa/classifications/entities/
entity_relationships collections; the reports collection's text field
is explicitly excluded from that grant. This dataset ingests AIID only
via its official weekly snapshot archive (docs/SOURCE_LICENSES.md
section 1.2a) and reads the licensed incidents.description field only as an
ephemeral, in-memory classification signal that is never persisted --
what is kept is the AIID-authored title (a short headline, not
narrative) plus structured, non-narrative facts (id, date, url, entity
slugs, taxonomy fields). As with AIAAIC, this reduction means
LICENSE-DATA carries no BY-SA carve-out for AIID. AIID's
content-licensing exposure is RESOLVED, not open, for the population
that ships -- the opposite posture from AIAAIC's, not the same one.
docs/audits/E23-aiid-marking-ruling-2026-07-30.md finds AIID's maker
(Responsible AI Collaborative, Inc.) is U.S.-situated -- principal
place of business, venue, and tax-exempt status all in the United
States -- and that situs closes both routes CC BY-SA offers to
share-alike: U.S. Copyright Office Circular 33 categorically excludes
titles and short phrases from copyright, so the retained AIID title
carries no protectable expression; and a U.S.-situated maker fails
both the UK and EU sui generis database-right qualification tests, so
no such right subsists in AIID's collections to begin with. This is
the opposite of AIAAIC's still-open UK-situated question, where UK/EU
law has no equivalent categorical carve-out -- AIID and AIAAIC are not
two postures toward the same question; they are two legal systems
answering it differently. Empirically, no row in the shipped corpus
carries a content_license marker for AIID under any measurable
definition of "an AIID row": 0 of 1,466, counting the union of the
AIID tag, the aiid_id field, the AIID-<n> source-ID, and the
incidentdatabase.ai reference-URL signals -- not merely the narrowest
tags-only count. This resolution covers only the ~1,463-row population
that ships via the sanctioned-snapshot path (section 1.2a); it does
not extend to the 63 hand-curated rows in ingest/aiid_incidents.json
or the 1,457 AIRI Navigator rows sharing AIID's source-ID convention,
which carry unreviewed, uncleared risk and are kept out of the shipped
corpus only by current merge-order mechanics, not by compliance
design -- if a future pipeline or merge-order change ever causes that
population's own text to reach data/incidents.json, this resolution
does not carry over, and that population reopens for its own review
before it ships. See docs/SOURCE_LICENSES.md section 1.2a and
docs/audits/E23-aiid-marking-ruling-2026-07-30.md for the full
reasoning.

The OECD AI Incidents and Hazards Monitor (AIM) is a third upstream
source carrying an active content obligation, alongside AIAAIC and
AIID above. AIM's own terms self-incorporate OECD's general
Data-reuse conditions -- "Your use of the OECD AI incidents and
hazards monitor (previously AI Incidents Monitor) ("AIM") is subject
to the terms and conditions found at www.oecd.org/termsandconditions"
-- which permit OECD content to be "extract[ed] from, download[ed],
cop[ied], adapt[ed], print[ed], distribute[d], share[d] and embed[ded]
... for any purpose, even for commercial use," conditioned on
attribution in the format OECD (year), (dataset name), (data source)
DOI or URL (accessed on (date)). That grant does not reach the
incident narrative: AIM's own methodology page states that any
copyrights, trademarks, or other intellectual property rights
included in AIM are the property of their respective owners, and that
"the OECD cannot and do not grant any rights to use or otherwise
exploit these protected materials included herein." This project does
not present OECD's general terms as covering that narrative text.

Per docs/audits/E21-oecd-narrative-licence-2026-07-30.md (E21,
2026-07-30), what AIM actually supplies for its title/summary/
harm_type/severity/affected-stakeholder/country metadata is not OECD
staff prose and not quoted news text: it is LLM-generated (OpenAI's
o3-mini) from the top three source articles of each event, selected
from different news outlets -- machine output derived from
copyrighted third-party news of unresolved ownership. Per that
finding, this project's ingest builds its description field from
structural facts and a source link only (the AIM source id, the
incident date, named entities, the classified attack vector, and the
AIM page link) -- never from summary or evidences -- mirroring the
pattern scripts/ingest_external.py::ingest_aiid_oecd_bridge() already
uses for AIID/OECD cross-listed entries. This reduction is
implemented and merged. Per-entry OECD attribution -- OECD (year), AI
Incidents and Hazards Monitor, url (accessed on date) -- is likewise
implemented and merged, reaching every OECD-AIM-sourced row plus a
further set of entries whose only OECD-AIM content was absorbed into
an AIID entry during merge but which still carry an oecd.ai page
reference via the AIID-OECD bridge path.

The title field is NOT resolved by this reduction, and this is not
silent: it still ships the same LLM-generated text verbatim on 3,667
of 3,829 OECD-AIM-sourced rows -- the other 162 were merged into an
entry whose own title, from a different upstream source, took
precedence instead. This is tracked as a separate, open question at
the same posture as AIAAIC's retained-headline question above
(E15/D17) -- a shorter label, not a multi-sentence digest, but not
dispositive of the underlying copyright question. Unlike AIAAIC, no
row-level content_license marker is emitted for OECD-derived rows
today; whether one is warranted is a separate, not-yet-scoped
question. Separately: OECD-AIM-sourced entries also carry verbatim
third-party news-article headlines in their references list (not
AIM's own text) -- this is cleared under E21 section 5.1 and decision
D17, not an open exposure, and is named here only so a reader of this
file alone knows it is there. See docs/SOURCE_LICENSES.md section 1.5
for the full framing and docs/audits/E21-oecd-narrative-licence-2026-07-30.md
for the full reasoning.

The full CC BY 4.0 Public License text is in LICENSE-DATA at the
repository root.
