Metadata-Version: 2.4
Name: cobolwork
Version: 0.2.76
Summary: COBOL, JCL and CICS security analysis: reference-format parser, cross-program data flow, mainframe credential rules. Runs on Node.js 18 or later.
Author-email: Portll <john@portll.net>
License-Expression: AGPL-3.0-or-later
Project-URL: Homepage, https://github.com/Portll/cobolwork
Project-URL: Source, https://github.com/Portll/cobolwork
Project-URL: Issues, https://github.com/Portll/cobolwork/issues
Keywords: cobol,jcl,cics,mainframe,sast,security
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Topic :: Security
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
License-File: LICENSING.md
License-File: THIRD-PARTY-NOTICES.md
Dynamic: license-file

# cobolwork

[![CI](https://github.com/Portll/cobolwork/actions/workflows/ci.yml/badge.svg)](https://github.com/Portll/cobolwork/actions/workflows/ci.yml)
[![Node](https://img.shields.io/badge/node-%E2%89%A518-informational)](https://nodejs.org)
[![Dependencies](https://img.shields.io/badge/dependencies-0-informational)](package.json)
[![Licence](https://img.shields.io/badge/licence-AGPL--3.0--or--later-informational)](LICENSE)

Cobolwork offers security analysis for COBOL, JCL and CICS, with no runtime dependencies. It reads
the source the way a compiler does and follows untrusted data across program boundaries.

It runs on its own. It can also run as a lane inside commitwork, Portll's CI and security runner.

## Why it exists

Nothing open source reads COBOL for security. Measured 2026-09-16: Semgrep, CodeQL, SonarQube
Community Edition and PMD ship no COBOL rules; gitleaks and TruffleHog have no rule for a RACF
password, a JCL `PASSWORD=` or a TSO logon; NIST's SARD holds zero COBOL cases. The one open-source
COBOL scanner, VisualCodeGrepper, matches regular expressions within a single file.

Meanwhile a Semgrep run over a COBOL-only repository loads three thousand rules, reads none of its
COBOL, exits zero and reports a clean result.

## Install

    npm install -g @portll/cobolwork
    pip install cobolwork

Either installs the command `cobolwork`; the PyPI package runs the same files with the Node.js on
your PATH. The same package is attached to each
[release](https://github.com/Portll/cobolwork/releases) as `cobolwork-<version>.tgz`, which is the
one to pin, and as `cobolwork.tgz`, which
`npm install -g https://github.com/Portll/cobolwork/releases/latest/download/cobolwork.tgz`
follows. To run from a checkout instead:

    git clone https://github.com/Portll/cobolwork && cd cobolwork && npm link

Node 18 or later. No dependencies, runtime or development. `npm link` puts `cobolwork` on your PATH;
adding `bin/` to PATH does the same thing. GnuCOBOL is needed only to regrade the parser or validate
benchmark cases.

## What it does

| Command | Information |
|---|---|
| `cobolwork scan <path>` | every rule set below, as JSON or SARIF |
| `cobolwork inventory <path>` | inventory (and what couldn't be read) |
| `cobolwork flow <path>` | where untrusted data reaches a sensitive operation, with the path it took |
| `cobolwork diff <repo> --base <ref>` | what a change reaches: layouts it moves in programs nobody edited, new call targets, findings it adds or removes |
| `cobolwork build <repo> [--base <ref>] [-- <compiler> …]` | the build gate: every finding ranked LOW to KNOWN-EXPLOITABLE, the build stopped on the ones the policy blocks and on compiler options that let a bad index corrupt storage, and the compiler run only on a pass |
| `cobolwork parse <file>` | one file's structure, for debugging |

## What it finds

| Area | What is reported |
|---|---|
| Data flow | untrusted input - the command line, a job's `PARM` or in-stream data, a CICS terminal or web request - reaching an OS command, dynamic SQL, a dynamic `CALL`, `LINK` or `XCTL`, a file name, the internal reader, a subscript or length, decimal arithmetic, a record key or a log, across programs |
| Checks | a check counts only where it runs first, on every route; one that leaves the value safe for the sink clears the route |
| Screen fields | a field the BMS map protects, read back and used to choose a record: the 3270 hidden form field |
| Exfiltration | database rows and file records leaving through web calls, sockets, MQ, extrapartition queues or service calls |
| CICS and CALL | a communication area read without `EIBCALEN`, a transfer to a variable program, a length longer than the area or the callee, a route around a sign-on |
| Privilege and logs | diagnostic transactions installed, command security off where it is used, credentials or personal data written to a log, input forging a log line, system error codes sent to a web client |
| JCL | credentials and security commands in in-stream data, destructive statements, `DLM=` tricks, FTP in cleartext or sending production data, production data touched by a test job |
| The source | names nothing declares (code that cannot compile), shadowed copybooks, payloads hidden in columns 73-80 or aimed at AI readers |
| The estate | production names outside production jobs, routable addresses, compiler and runtime versions with published advisories |

[docs/rule-sets.md](docs/rule-sets.md) describes each in full, with what it deliberately leaves out.
Vendor packs for CA ACF2 and Top Secret, Control-M and Connect:Direct load only for estates that
name them, and `rules/gitleaks-mainframe.toml` gives gitleaks the mainframe credential shapes it
lacks.

## What a finding looks like

The flaw is in neither file on its own: one program reads the command line and hands it on, the
other runs what it was handed. Both commands below run against this repository, so every number
here can be checked.

    $ cobolwork flow test/fixtures/dataflow

```json
{
  "rule": "argv-or-env-to-os-command",
  "sev": "crit",
  "cwe": "CWE-78",
  "path": "pos/P2.cbl",
  "line": 10,
  "program": "P2",
  "crossProgram": true,
  "hops": 4,
  "detail": "Command-line or environment input reaches an operating-system command routine: ACCEPT ... FROM COMMAND-LINE at pos/P1.cbl:8 reaches CALL 'SYSTEM' USING WS-LOCAL",
  "trace": [
    { "program": "P1", "item": "WS-IN",    "file": "pos/P1.cbl", "via": "source" },
    { "program": "P1", "item": "WS-CMD",   "file": "pos/P1.cbl", "via": "MOVE at pos/P1.cbl:9" },
    { "program": "P2", "item": "LK-CMD",   "file": "pos/P2.cbl", "via": "CALL 'P2' argument 1 at pos/P1.cbl:10" },
    { "program": "P2", "item": "WS-LOCAL", "file": "pos/P2.cbl", "via": "MOVE at pos/P2.cbl:9" }
  ],
  "related": [{ "path": "pos/P1.cbl", "line": 8, "detail": "ACCEPT ... FROM COMMAND-LINE" }],
  "sources": 1,
  "evidence": "path",
  "fingerprint": "45e55b502c397a499cc29212c0c28329"
}
```

`crossProgram` marks a path that left the file it started in. Findings carry no source text, so a
report can be stored and passed on without carrying the code with it.

`fingerprint` is what the finding is, rather than where it is printed today: the rule, the program
and the paragraph or section it sits in (the job, step and DD for JCL), and the flagged statement's
own text. No line number goes into it, so code added above a finding does not change it. `diff`
compares findings by it, and SARIF carries it as `partialFingerprints["cobolwork/v1"]`. Two findings
that only their position tells apart share one, and `summary.identity.shared` counts them.

### Findings and Claim Severity

Severity says how urgent a finding is. `evidence` says what the tool actually established, which
decides who acts on it. Every rule declares one of eight finding types:

| `evidence` | What the finding claims | Who acts |
|---|---|---|
| `path` | untrusted input was traced to a sensitive operation, and `trace` is the route | the program's owner |
| `construct` | the construct is a defect wherever it sits; no input has to reach it | the owner, or whoever can rotate a credential |
| `tampering` | the source is arranged so a reader or resolver sees something other than what runs | a reviewer, before merge |
| `advisory` | a pinned compiler or runtime matches a published advisory | whoever owns the build |
| `exposure` | information about the estate is written into source | the owner |
| `change` | a change moves an interface or adds a call target (`diff` only) | the reviewer of that change |
| `coverage` | the analysis stopped following here | nobody's code; read more, or accept the limit |
| `context` | describes the estate (an entry point, a product in use) | nobody; it asserts no defect |

`coverage` and `context` are exactly the `info` rules (not defects). A consumer that counts these
findings should leave them out of its counts.

None of the kinds says *exploitable*, on purpose. Whether a route can be used also depends on who may
start the transaction or job that reaches it, and on what the running system enforces, and no
repository holds those facts. A finding names the entry points that reach it (`startedBy`). `path`
is the strongest claim the tool makes: a route found by reading the code, not by running it. Its
precision has been measured on benchmark cases this project wrote, not yet on an independently
labelled corpus.

### What each finding lets someone do, and the fix

A defect finding carries two more facts, the same for every finding of its rule: what someone can do
once the route or construct is present, and the standard fix. They are what makes a `path` or
`construct` finding read as a hole to act on rather than a location. A report carries them once per
rule, keyed by rule id beside `ruleText` and `ruleCwe`, so a finding does not repeat them. For the
finding above:

```json
"ruleImpact": { "argv-or-env-to-os-command": "Whoever sets the program's command line or environment runs an arbitrary operating-system command with the program's authority" },
"ruleRemedy": { "argv-or-env-to-os-command": "Build the command only from fixed literals; never place input in the argument of CALL 'SYSTEM' or BPXWDYN. If it must vary, choose from an allow-list of known commands" }
```

The `who` is the finding's own — the entry points it carries in `startedBy` — and the impact
completes it. SARIF puts the impact in each rule's `fullDescription` and the fix in its `help`, which
is where GitHub code scanning shows a recommendation. The `info` kinds carry neither, because a fix
would assert a defect they do not claim; a site-gated rule carries them once the estate's fact turns
it into a defect, and not before.

### What it could not read

A scan says what each rule set read, not only what it found:

    $ cobolwork scan bench/cases --quiet

```json
"summary": {
  "findings": 127,
  "bySeverity": { "high": 28, "crit": 16, "med": 16, "low": 9, "info": 58 },
  "byEvidence": { "tampering": 6, "path": 31, "advisory": 1, "construct": 31, "coverage": 6, "context": 52 },
  "bySet": {
    "inventory": { "filesScanned":  83, "filesUnreadable": 0 },
    "flow":      { "filesScanned":  83, "filesUnreadable": 0 },
    "cics":      { "filesScanned":  42, "filesUnreadable": 0 },
    "hidden":    { "filesScanned": 116, "filesUnreadable": 0 },
    "copybook":  { "filesScanned":  88, "filesUnreadable": 0 },
    "jcl":       { "filesScanned":  28, "filesUnreadable": 0 },
    "build":     { "filesScanned":   2, "filesUnreadable": 0 },
    "recon":     { "filesScanned": 116, "filesUnreadable": 0 },
    "vendor":    { "filesScanned":   0, "filesUnreadable": 0 },
    "opaque":    { "filesScanned":  83, "filesUnreadable": 0 },
    "web":       { "filesScanned":   9, "filesUnreadable": 0 },
    "compile":   { "filesScanned":  83, "filesUnreadable": 0 },
    "priv":      { "filesScanned": 234, "filesUnreadable": 0 },
    "log":       { "filesScanned":  13, "filesUnreadable": 0 }
  },
  "coverageIncomplete": false,
  "setsIncomplete": [
    { "set": "flow", "kind": "configuration",
      "why": "6 EXEC CICS WRITEQ TD statement(s) write to an extrapartition queue, and nothing says whether it reaches the internal reader: name the region's INTRDR DDs as internalReaderDds (or the queues as internalReaderQueues) in cobolwork.site.json" },
    { "set": "jcl", "kind": "configuration",
      "why": "5 FTP transfer(s) send a named dataset, and cobolwork.site.json names no production qualifier, so the production-data rule did not run on them" },
    { "set": "recon", "kind": "configuration",
      "why": "no cobolwork.site.json: the production-name and production-dataset rules did not run, because nothing declares what production means in this estate" }
  ],
  "flowModel": "byte-range",
  "toolVersion": "0.2.0"
}
```

`coverageIncomplete` is the field to read first. An unresolved copybook, an unreadable file or a
symlink leading out of the tree sets it, because a finding count over source nobody read is not a
clean result.

Some rules need facts no repository holds: which dataset qualifiers are production, which DDs reach
the internal reader, the compiler options and runtime versions in use, which libraries are
authorised. They go in `cobolwork.site.json`, and `node diag/propose-site.mjs <path>` drafts one from
the estate's own JCL for a person to correct. Without a fact, the rule that needs it says it did not
run, under `setsIncomplete`, rather than reporting a clean result. The benchmark tree has no site
file, which is why three sets say so above. `advisoryCoverage` names the products the advisory rules
searched, so a scan with no advisory finding says what that silence covers.

### Findings someone has already judged

    $ cobolwork baseline . --reason "replaced by a fixed command table in Q1" --who jsmith --expires 2027-03-31

writes `cobolwork.baseline.json`, one entry per finding keyed by fingerprint: `accept`,
`false-positive` or `wont-fix`, with the reason, who, and until when. A later scan moves what it
covers into `suppressed`. Every suppression expires, and the finding comes back when it does. A
baseline inside the scanned tree cannot hide tampering - a hidden payload, or an instruction aimed
at an AI reader - so that takes `--baseline <file>` from outside the tree. `--no-baseline` applies
none.

## Stopping a build

`cobolwork build` is a CI step, and its exit status is its verdict: 0 pass, 1 fail, 3 undecided, 4 the
compiler failed after a pass, 2 it could not run. No model is asked anything; every check is the
engine's reading of the tree.

    cobolwork build . --base origin/main -- cobc -x -o payroll PAYROLL.cbl

By default HIGH, CRIT and KNOWN-EXPLOITABLE findings block, and so does any finding, at any tier,
that would let its author escalate privilege or change data: input choosing a command, a program, an
SQL statement or a job, or choosing which record is rewritten or where in storage a write lands. MED
and LOW findings outside those two classes are reported and do not block. With `--base`, a finding
blocks only if the change introduced it, except CRIT and KNOWN-EXPLOITABLE, which block wherever they
are; the change is judged by the policy, waivers and site file of its base, so it cannot relax its own
gate. An incomplete scan is undecided, never a pass.

The compiler options are held to the policy too. A program compiled without `SSRANGE`, or with
`SSRANGE(MSG)`, which reports a bad subscript and carries on, fails; for `cobc` the missing `-fec`
checks are added to the command. `cobolwork.policy.json` changes any of this, and `--policy <file>`
names an organisation's floor, which a repository can tighten and never loosen.
[docs/spec/build-gate.md](docs/spec/build-gate.md) is the full specification.

## How accurate it is

The parser is graded against GnuCOBOL's own listing (`cobc -t -Xref -ftsymbols`), which reports every
data item with the size the compiler computed, every label, called programs, and which references
write to a field. `diag/grade-against-gnucobol.mjs` runs that comparison over a corpus, on
repositories never used while building the parser. The 100 and 300 sets were measured 2026-09-18,
the 500 set 2026-09-26:

| Corpus | Files the compiler accepted | Data items | Sizes | Labels and calls |
|---|---|---|---|---|
| 100 repositories, held out | 489 | 100% recall, 100% precision | 0 disagree of 15,888 | 100% |
| 300 repositories, held out | 2,210 | 99.9% / 100% | 8 disagree of 94,786 | 100% / 99.7% |
| 500 repositories, held out | 21,624 | 99.98% / 99.998% | 5,881 disagree of 733,835 | 100% / 100% |

On the 500 set, 5,643 of the 5,881 size disagreements come from one repository that vendors a COBOL
research dataset; the other repositories disagree on 238 of 343,444. The grade covers only programs
GnuCOBOL accepts, so a program with EXEC SQL or EXEC CICS is graded only through the precompiler
stand-in in `diag/precompiler.mjs`, which rewrites what the parser would otherwise have to read. The
tests compare the parser with the compiler's answers kept in `test/fixtures/parser/*.golden.json`,
so they run without GnuCOBOL.

`bench/cases/` holds 93 CWE-labelled cases, each paired with a near-miss negative: the same shape
with the flaw removed. `node bench/run.mjs` scores any scanner's findings against them, by rule and
file, never by line, and `npm test` fails if any case scores differently from its declaration.
`--validate` compiles every COBOL case with GnuCOBOL and checks every JCL case against the
statement grammar, a weaker witness, and says so.

## Coverage on a busy machine

A scan stops before it exhausts memory, reports how much of the tree it read, and sets
`coverageIncomplete`. **Read that before the finding count.** On one 4,086-file repository a starved
run reported 26 findings and a clean one 2,890.

The number it watches is the lesser of the heap's headroom and the machine's free memory, and the
second is the whole machine. On a build agent running other jobs, or in a container whose limit is
smaller than the host, that means a scan's coverage is decided by what else is running. Say what your
share is and it will use that instead:

```sh
COBOLWORK_FREE_MEMORY_MB=4096 cobolwork .
```

It is a statement, not a limit: the heap is still watched, so an over-generous number does not turn
the guard off, it just stops the host's load from deciding. A spawned scan inherits this property.

## Compliance

Every rule is mapped to the clause of each framework that makes it an obligation, quoted verbatim:

| File | Instrument | Rules mapped |
|---|---|---|
| `rules/compliance-dora.json` | Regulation (EU) 2022/2554 (DORA) | 156 |
| `rules/compliance-ffiec.json` | FFIEC IT Examination Handbook | 154, and 2 recorded as unmapped |
| `rules/compliance-nist80053.json` | NIST SP 800-53 Rev. 5.2.0 | 154, and 2 recorded as unmapped |

`node diag/map-compliance.mjs` refuses to write a quote the cached instrument does not contain, and
every scan carries the mapping as `ruleCompliance`. The clause choice is a judgement that no
qualified assessor has reviewed. Where no control genuinely covers a rule, as for committing an LPAR
name to a repository, the rule is recorded as unmapped with a reason rather than mapped to the
nearest control that reads plausibly. What else the mapping does not claim is in
[docs/rule-sets.md](docs/rule-sets.md#compliance).

## Tests

    npm test

Tests that need a tool which is absent record a skip naming it. A skipped check is not a passing one.
Open work is in [BACKLOG.md](BACKLOG.md).

## Security

Vulnerabilities in cobolwork itself go through [SECURITY.md](SECURITY.md), privately. A false
positive or a missed finding is an ordinary issue, and a wanted one: precision and recall are
measured and published here, so a report that moves either is the most useful thing you can send.

## Licence

AGPL-3.0-or-later. See [LICENSE](LICENSE) and [NOTICE](NOTICE). That covers this project's own work; material belonging to
others — AWS CardDemo fixtures, IBM interface layouts, quoted regulatory clauses — is listed in
[THIRD-PARTY-NOTICES.md](THIRD-PARTY-NOTICES.md) with its own licence.

If your organisation's policy refuses AGPL, [LICENSING.md](LICENSING.md) says what else is available
and what has to be true first. Running the scanner over your own source triggers nothing in the AGPL
that unmodified internal use does not already satisfy; that page explains why, which is often the
whole of the objection.
