Metadata-Version: 2.4
Name: profgen
Version: 0.0.3
Summary: Convert candidate CVs into a standardised Word profile, with no invented facts.
Author-email: Kevin Steptoe <kevin.steptoe@gmail.com>
License: MIT
Project-URL: Homepage, https://github.com/ksteptoe/profgen
Project-URL: Repository, https://github.com/ksteptoe/profgen
Project-URL: Documentation, https://profgen.readthedocs.io/
Keywords: cv,resume,docx,anthropic,claude,semiconductor,recruitment
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Other Audience
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Office/Business
Classifier: Topic :: Text Processing :: Markup
Requires-Python: >=3.13
Description-Content-Type: text/markdown
License-File: LICENSE.txt
License-File: AUTHORS.md
Requires-Dist: click>=8.1
Requires-Dist: pydantic>=2.6
Requires-Dist: python-docx>=1.1
Requires-Dist: pdfplumber>=0.11
Requires-Dist: anthropic>=0.40
Provides-Extra: pdf-fast
Requires-Dist: pymupdf>=1.24; extra == "pdf-fast"
Provides-Extra: pdf-word
Requires-Dist: comtypes>=1.4; sys_platform == "win32" and extra == "pdf-word"
Provides-Extra: docs
Requires-Dist: sphinx>=7; extra == "docs"
Requires-Dist: myst-parser>=2; extra == "docs"
Provides-Extra: dev
Requires-Dist: profgen[docs,pdf-fast,pdf-word]; extra == "dev"
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: pytest-cov>=5; extra == "dev"
Requires-Dist: pytest-xdist>=3.6; extra == "dev"
Requires-Dist: pytest-timeout>=2.3; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: mypy>=1.8; extra == "dev"
Requires-Dist: reportlab>=4.0; extra == "dev"
Requires-Dist: pre-commit>=3.7; extra == "dev"
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: twine>=5; extra == "dev"
Dynamic: license-file

# profgen (profile generator)

[![Project generated with PyScaffold](https://img.shields.io/badge/-PyScaffold-005CA0?logo=pyscaffold)](https://pyscaffold.org/)

> Convert candidate CVs into standardised Word profiles — without inventing facts.

`profgen` (the **profile generator**) turns a candidate CV
(PDF/DOCX/TXT) into a standardised Word profile through a
verbatim-extract → typed-structure → grounding-check → render → review-report
pipeline. The profile is rendered against a template you supply, so any house
style — including a private or corporate one — can be applied without the
template living in the package.

## The one hard rule: no invented facts

Omitting information is acceptable; fabricating a company, tool, date,
qualification, institution or project is a defect. Concretely:

- Anything absent from the source CV is marked `"Not stated"` (scalars) or left
  as an empty list — never guessed.
- A deterministic, LLM-independent **grounding check** verifies that every
  extracted **hard fact** — company, tool, certification, institution, dates and
  project domain — actually appears in the source text. Anything it cannot find is
  flagged in the review report. Inferred or curated content (the summary, the
  derived grade, the applied discipline) is exempt from grounding and surfaced in
  the review report for human sign-off instead.
- **Employers are anonymised.** Experience is rendered as `Project N | <domain>`
  rather than by company name (the company is still extracted, purely so the
  grounding check can confirm nothing was invented).
- **No fabricated hard facts.** The renderer derives nothing itself: a total
  years-of-experience figure is never computed from dates and seniority titles are
  never invented. On the Claude path the skills table's per-skill "Years
  Experience" column may be *estimated* from the supporting employment span
  (`balanced`/`bold` only) — but every such estimate is recorded as an inferred
  value in the review report for sign-off; otherwise the column reads
  `"Not stated"`.

Each conversion therefore writes two files: the `.docx` profile **and** a sibling
`*.review.md` listing missing information and everything to verify before customer
submission.

## Installation

`profgen` is published on [PyPI](https://pypi.org/project/profgen/) (Python 3.11+):

```bash
pip install profgen
```

Or install from source for development:

```bash
git clone https://github.com/ksteptoe/profgen
cd profgen
make dev                  # editable install with all dev/docs extras
# or, equivalently:
pip install -e ".[dev]"
```

## Quickstart

```bash
# 1. Generate a starter .docx style-donor template (neutral default styles).
profgen make-template templates/profile_template.docx

# 2. Convert a CV offline (no API key, no network) — produces out.docx AND out.review.md.
profgen convert cv.pdf --output out.docx --offline
```

When `--output` is omitted the profile is written to `<identifier>_profile.docx`
in the current directory — named from the **abstracted identifier** (e.g.
`VV_profile.docx`), resolved after structuring so the candidate's real name never
leaks into the output filename — with the review report alongside. (The name falls
back to the source stem only when no name could be abstracted.)

To convert a whole folder of CVs at once, pass a directory with `-R/--recursive`:

```bash
profgen convert ./cvs -R --offline
```

This recursively finds every `.pdf`/`.docx`/`.txt` CV beneath the directory,
converts each one, and writes each profile (and its `*.review.md`) **next to its
own source file**, finishing with a `Converted N profile(s).` summary. Any existing
target is backed up to a deterministic `<name>.bak` (then `.bak.1`, `.bak.2`, …)
before being overwritten, and generated artefacts and hidden files are skipped.
`--output` cannot be combined with `-R` (each profile goes next to its source).

To collect every profile in **one folder** instead, add `--output-dir DIR`. It
redirects the derived `<identifier>_profile.docx` (and its `*.review.md`) into `DIR`
— the directory only; the filename stays the PII-safe abstracted identifier — and
creates `DIR` if it does not exist. Unlike `--output`, `--output-dir` works with
`-R`: it flattens the whole tree into the one directory, backing up to `<name>.bak`
on an identifier collision. `--output-dir` and `--output` cannot be combined.

Add `-c/--cost` to print the run's estimated Claude API cost once the profile is
written — the input/output token counts and the estimate **headlined in GBP (£)**.
Anthropic bills in USD, so the figure is computed from the published
`claude-sonnet-4-6` pricing ($3.00 per million input and $15.00 per million output
tokens) and converted USD→GBP at a default rate of `0.79`, with the USD basis and
the rate shown so it stays verifiable. Override the rate with `--gbp-rate FLOAT`
(e.g. `--gbp-rate 0.80`) — FX moves, so set it to your own current rate when the
figure matters. On `--offline` it reports that no API call was made, and on a
recursive run it adds a per-file line plus a grand total.

Add `--pdf` (or set `[render] pdf = true` in a `--config` file) to also write a
sibling **PDF** next to the `.docx` profile (e.g. `VV_profile.docx` →
`VV_profile.pdf`). It works with `--output`, `--output-dir` and `-R/--recursive`.
PDF output **requires LibreOffice** — the conversion shells out to the headless
`soffice` binary (it must be on `PATH`), since a faithful DOCX→PDF that keeps the
branded header, footer, logo and styles needs a real office renderer. If
LibreOffice is not found, `convert` fails fast with a clear message before doing
any work (install LibreOffice — Debian/Ubuntu `sudo apt-get install
libreoffice-writer`, macOS `brew install --cask libreoffice`, Windows install
LibreOffice and add `soffice` to `PATH` — or drop `--pdf`).

`cv-formatter` is an identical alias for `profgen`, and `python -m profgen` works
too. Run `profgen convert --help` for the full option list.

### Bring your own template

The renderer binds content to ten **logical roles** — `title`, `section_heading`,
`date_heading`, `project_method`, `subsection_label`, `body`, `bullet`,
`domain_line`, `legal` and `table` — rather than to fixed style names. By default
each role maps to a neutral built-in or starter style (`DEFAULT_STYLE_MAP`), e.g.
`title` → `Profile Title`, `date_heading` → `Profile Date`, `body` → `Normal`,
`bullet` → `List Bullet`, `legal` → `Profile Legal`, `table` → `Table Grid`. The
full role-to-style table is in the
[user guide](https://profgen.readthedocs.io/en/latest/userguide.html#logical-roles-not-fixed-style-names).

To apply your own house style, pass your branded document with `--template` and a
TOML **style map** with `--style-map` that points each role at the real paragraph
style names in *your* document:

```bash
profgen convert cv.pdf --template my_template.docx --style-map my-style-map.toml
```

```toml
# my-style-map.toml — map the logical roles to YOUR template's style names.
title        = "My Heading Style"
date_heading = "My Date Style"
legal        = "My Legal Style"
```

The map may be partial: any role you omit falls back to its default (so a legacy
map naming only the original five roles still loads). This is how a private or
corporate template can be applied without it ever living in the package. The donor
document's header, footer, fonts and logo media are preserved. The complete,
ready-to-copy map ships at `examples/style-map.example.toml`. For how to get the
full branded look — and the warning `convert` prints when the donor carries no
header, footer or logo — see
[Branding your profiles](https://profgen.readthedocs.io/en/latest/userguide.html#branding-your-profiles)
in the user guide.

### One-step branded profile (`make profile`)

`make profile` renders a finished, branded profile from a candidate's CV using
your private (gitignored) `local/` assets. It is the one-command wrapper around
`profgen convert`.

```bash
cp examples/style-map.example.toml local/style-map.toml   # required — then edit
# drop your branded template at local/template.docx        # required
make profile CV=cv.pdf                                     # Claude path (needs ANTHROPIC_API_KEY)
make profile CV=cv.pdf OFFLINE=1                           # deterministic, network-free path
make profile CV=cv.pdf OUT=out.docx CREATIVITY=bold DISCIPLINE="Verification"
```

**What it does**

1. **Checks prerequisites** — fails early with a clear message if `CV=` is
   missing, or if `local/template.docx` / `local/style-map.toml` are absent.
2. **Runs the pipeline** — `profgen convert <CV>` against your branded template
   and style map, writing the `.docx` profile and its sibling `*.review.md`.

It expands to:

```bash
profgen convert "$CV" \
  --template local/template.docx --style-map local/style-map.toml \
  --creativity balanced \
  [--grade-map local/grade-map.toml]          # auto-added when the file exists
  [--discipline-config local/discipline.toml] # auto-added when the file exists
  [--job-spec local/job-spec.txt]             # auto-added when the file exists
  [--discipline "..."] [--output ...] [--offline]
```

**Inputs** (all under the gitignored `local/`):

| Input                   | Role                                                                          | Required?            |
|-------------------------|-------------------------------------------------------------------------------|----------------------|
| `local/template.docx`   | your branded style-donor (header/footer/logo + named styles)                  | **yes**              |
| `local/style-map.toml`  | maps logical roles → your template's real style names (copy `examples/style-map.example.toml`) | **yes**              |
| `local/grade-map.toml`  | years-of-experience → grade ladder (copy `examples/grade-map.example.toml`)    | optional, auto-wired |
| `local/discipline.toml` | per-discipline priority skills + tool groups (copy `examples/discipline.example.toml`) | optional, auto-wired |
| `local/job-spec.txt`    | per-run brief that augments the discipline defaults                            | optional, auto-wired |

The optional three are added **only when the file exists**, so a minimal setup
with just the template and style map works.

**Variables**

- `CV=` — the source CV (`.pdf` / `.docx` / `.txt`). Required.
- `OUT=` — output path (default `<identifier>_profile.docx` in the current directory).
- `OFFLINE=1` — use the deterministic, network-free heuristic instead of Claude.
- `CREATIVITY=` — `strict` / `balanced` (default) / `bold`; gates only the inferred fields.
- `DISCIPLINE=` — overrides the inferred discipline.

**Default vs offline**

- **Default** — the Claude path; needs `ANTHROPIC_API_KEY`. Produces the curated,
  client-grade profile.
- **`OFFLINE=1`** — no network, fully deterministic; a degraded skeleton (grounded
  facts only, no curation) that **ignores** `CREATIVITY`, `DISCIPLINE` and the
  job-spec. Useful for a quick structural check without an API key.

The `local/` directory and `.env` are **gitignored**, so confidential templates,
configs and API keys stay out of the repository. Override the asset directory
with `LOCAL_DIR=` (e.g. `make profile CV=cv.pdf LOCAL_DIR=clientA`) to keep
per-client asset sets.

### Targeting and identity

Two further TOML files let you describe your internal grade ladder and the
disciplines you staff. Starter files ship under `examples/`; copy each to the
gitignored `local/` and edit it, then pass it on the command line:

```bash
cp examples/grade-map.example.toml   local/grade-map.toml
cp examples/discipline.example.toml  local/discipline.toml
profgen convert cv.pdf --grade-map local/grade-map.toml \
    --discipline-config local/discipline.toml --discipline Verification
```

- **`grade-map.toml`** — your years-of-experience → grade-title ladder.
- **`discipline.toml`** — per-discipline priority skills and tool groupings.

All of `--grade-map`, `--discipline`, `--discipline-config`, `--id-mode`, `--id`,
`--creativity` and `--job-spec` are **active**. The grade is derived from the
candidate's stated years (and stays `"Not stated"` when no years are given), the
operator's discipline is applied, and the candidate's identity is abstracted —
`--id-mode initials` renders "Valentin Vasilescu" as "VV", `--id-mode number`
gives a stable number, `--id` forces an explicit value, and `--id-mode none` keeps
the real name. The real name is always retained for grounding but only shown with
`--id-mode none`.

On the **Claude path**, `--creativity {strict,balanced,bold}` (default `balanced`)
tunes how much interpretive licence the structuring step takes over *inferred*
fields — hard facts stay strictly grounded at every level — and `--job-spec FILE`
plus a matched `--discipline` steer skill prioritisation and the summary toward
that role. The `--offline` heuristic client ignores these (it curates nothing).
See the
[user guide](https://profgen.readthedocs.io/en/latest/userguide.html#configuration-files)
for the full details.

### One config file for everything (`--config`)

Rather than pass each input as a separate flag, you can gather the style map, grade
ladder, discipline config, job-spec brief and render defaults into a single unified
TOML file and point `convert` at it:

```bash
profgen convert cv.pdf --config local/config.toml
```

Explicit flags still override the file (flag → `--config` value → built-in
default), so every flag-driven invocation keeps working unchanged. A starter ships
at `examples/config.example.toml`; `make profile` auto-wires `local/config.toml`
when present.

The unified file also drives the profile's **section order and selection**: a
`[[section]]` list chooses which body sections render and in what order (with an
optional per-section heading override), and can switch on the off-by-default
`certifications` and `languages` sections. The title and legal footer stay pinned
first and last, and `[render].legal_text` overrides the closing footer wording. See
the
[user guide](https://profgen.readthedocs.io/en/latest/userguide.html#the-unified-configuration-file)
for the full schema.

## Offline vs real Claude path

The structuring stage has two interchangeable backends behind one interface:

- **Offline (`--offline`)** — the deterministic, network-free
  `HeuristicStructuringClient`. Needs no API key, makes no network call, and is
  what the entire test suite uses. Ideal for plumbing checks and CI.
- **Real Claude (default)** — the `ClaudeStructuringClient`, which calls the
  Anthropic API and needs `ANTHROPIC_API_KEY`. This path is deliberately **never**
  exercised in CI; it is smoke-tested only behind an explicit opt-in (see
  `examples/smoke_real_path.py`).

## Example

A runnable, fully-offline example builds a profile from a bundled synthetic CV
with no API key:

```bash
.venv/bin/python examples/build_example_profile.py
```

It reads `examples/input_cvs/sample_cv.txt`, runs the offline pipeline, and writes
the profile and its review report into `examples/output_profiles/` (gitignored).

## Development

```bash
make dev      # editable install with all dependencies
make test     # run the fully-offline test suite
make lint     # ruff
make format   # ruff --fix
make docs     # build the Sphinx HTML User Guide
make docs-pdf # build a single PDF of the docs (needs a LaTeX toolchain)
```

Quality gates: `ruff` clean, `mypy --strict` clean (scoped to `src/`), and
`pytest` green with the network disabled. The Sphinx **User Guide** (`make docs`,
or `docs/userguide.md`) is the full pipeline walkthrough and the shipped reference
for every option and configuration file.

`make docs-pdf` produces `docs/_build/latex/profgen.pdf`. It needs a system LaTeX
toolchain on `PATH` — `xelatex`, `latexmk`, and `makeindex` (install
[TeX Live](https://www.tug.org/texlive/) or, on Windows,
[MiKTeX](https://miktex.org/)). The toolchain is **not** pip-installable and is
optional: the target fails fast with a clear message if `latexmk` is missing.

## Note

This project has been set up using [PyScaffold](https://pyscaffold.org/) 4.6
with the [ClickStart](https://github.com/ksteptoe/pyscaffoldext-ClickStart) extension.
