Metadata-Version: 2.4
Name: prompt-tailor
Version: 0.2.0
Summary: Model-aware prompt refinement tool for Claude Code users
License: MIT License
        
        Copyright (c) 2026 Createyouracccount
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Repository, https://github.com/Createyouracccount/PromptTailor
Project-URL: Issues, https://github.com/Createyouracccount/PromptTailor/issues
Keywords: claude,claude-code,prompt,prompt-engineering,mcp
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# PromptTailor

> Model-aware prompt rewriting for Claude Code. Write a rough request — get it rewritten the way your current Claude model works best.

[한국어 README](README.ko.md)

## Why

Claude Fable 5, Opus 5, Sonnet 5, and Haiku respond best to *different* prompt styles — Fable wants goals and constraints in prose (no step lists), Opus over-verifies if you tell it to double-check, Haiku wants small numbered steps. PromptTailor keeps these differences as data ([model profiles](prompt_tailor/profiles/)), detects which model you're running, and rewrites your rough request to match — also routing by task intent (fix / build / research / refactor / docs).

Your input language is preserved: English in → English out, Korean in → Korean out.

## Example

```
$ prompt-tailor "fix the login bug asap, users keep getting logged out" --model fable-5
```

> Users are repeatedly logging out unexpectedly. Before fixing, investigate: exact
> reproduction steps (when and under what conditions does this happen?), when this
> started, relevant error logs or console messages, and the login/session management
> code structure.
>
> Once you've identified the reproduction path and root cause, apply the minimum fix
> to prevent unintended logouts. Scope: session and login logic only — do not modify
> other features.
>
> Validation: confirm the issue no longer reproduces through direct testing, or verify
> that related tests pass.

Notice what happened: vague urgency ("asap") became an investigation directive, a scope boundary, and a validation criterion — and nothing was invented. Unknowns become investigation steps; any added specifics are tagged as assumptions.

## Install

**As a Claude Code plugin (recommended):**

```
/plugin marketplace add Createyouracccount/PromptTailor
/plugin install prompt-tailor@prompt-tailor
```

This gives you the `/pm` command with no path setup.

**As a CLI / MCP server:**

```bash
pip install prompt-tailor    # installs `prompt-tailor` and `prompt-tailor-mcp`
```

Requirements: Python 3.10+, the `claude` CLI installed and logged in (no separate API key). Verified on macOS/Linux; Windows untested.

## Usage

```bash
prompt-tailor "rough request" --model fable-5    # rewrite for a target model
prompt-tailor "rough request" --json             # JSON output
prompt-tailor "rough request" --concise          # faster, condensed meta-prompt
```

**Inside Claude Code** — `/pm rough request`: rewrites for your session's detected model, shows a one-line change summary, then executes the rewritten request. In auto mode, add the permission rule printed by `claude-code/install.sh` so prompts containing risky-looking words (e.g. "docker prune") aren't false-positive blocked — the backend only rewrites text.

**Hook auto mode (opt-in)** — rewrite every prompt automatically via a `UserPromptSubmit` hook. Run `bash claude-code/install.sh` for the settings snippet. Escape hatch: include `#raw` in a prompt to pass it through untouched. Prompts under 6 tokens or over 800 chars are skipped; if a rewrite doesn't finish within 28s it fails open (your original prompt goes through).

**Cursor / any MCP client** — a built-in stdio MCP server exposes `refine_prompt(raw, target_model, concise)`:

```jsonc
// ~/.cursor/mcp.json
{ "mcpServers": { "prompt-tailor": { "command": "prompt-tailor-mcp" } } }
```

```bash
claude mcp add prompt-tailor -- prompt-tailor-mcp   # register in Claude Code
```

## Same input, different models (real outputs)

Input: `로그인 버그 고쳐줘` ("fix the login bug" — Korean in, Korean out):

| target `fable-5` | target `haiku-4-5` |
|---|---|
| Prose: symptoms to identify first, then "find the cause and apply the simplest fix; verify by test or manual check". No step lists. | **Step 1** locate the bug (error message? where?) → **Step 2** fix ([assumption] tagged) → **Step 3** test with valid/invalid credentials → deliverable: fixed code + one-line commit message. |

Full texts in [eval/results.json](eval/results.json).

## Measured cost (n=2, 2026-08-15)

| What you pay | Measured |
|---|---|
| Per rewrite call (haiku) | **≈$0.03 API-equivalent** · ~1.8k output tokens · 18–33s wall. ~29.5k input tokens, but ~99% is `claude -p`'s own system prompt (cached: ~8k cache-write + ~21.6k cache-read); the meta-prompt itself adds only hundreds |
| Hook context injection | **+527 input tokens** in your main conversation (measured as token delta), and it stays in history for the rest of the session |
| Subscription users | No per-call bill — it consumes usage quota instead |

Raw data: [runs/cost_measurement.json](runs/cost_measurement.json), method: [eval/measure_cost.py](eval/measure_cost.py).

## Evidence — including the negative result

Every claim is backed by ledgered experiments (blind pairwise LLM judging; [LOOP_LOG.md](LOOP_LOG.md)):

- **Prompt quality** (golden set of 20 *vague* requests): 20/20 judged better than the original (clarity 5.0, fidelity 4.8, actionability 5.0) — [EVAL.md](EVAL.md)
- Model profiles produce structurally different rewrites: 5/5; intent routing beat profile-only meta 4–1–1
- **Task outcome pilot (n=3): the raw prompt won 3–0.** On *already-clear, self-contained* codegen tasks run headless, rewriting hurt: it inflated scope, its investigation directives stalled a run, and its verification demands made the executor fabricate test results — [eval/ab_task_outcome_results.json](eval/ab_task_outcome_results.json)

**What this means**: the measured benefit is on vague, underspecified requests — the golden set's territory. Already-clear requests should be left alone, so since v0.2.0 the rewriter runs a **clarity gate** first: if your request is already specific it returns it untouched (`action: keep`; the hook then injects nothing). Gate accuracy on a balanced 40-prompt benchmark: **39/40 (98%)** — vague recall 20/20, clear recall 19/20. Dataset, runner, and the one miss are documented in [BENCHMARK.md](BENCHMARK.md).

## What we do NOT guarantee

- **Higher task success is not proven.** Prompt-quality wins are judge-based; the only outcome data so far is the 3-task pilot above, which the raw prompt won.
- The rewriter occasionally adds specifics without an `[assumption]` tag (fidelity 4.8, not 5.0), and can misclassify intent.
- All experiments are small-n, single-LLM-judge, and were run in this repo's environment.

## Development

```bash
python3 -m unittest discover tests   # 37 offline tests, no LLM calls
python3 eval/run_eval.py             # golden-set evaluation (spawns claude)
```

Project docs (Korean): [PLAN.md](PLAN.md) · [ARCHITECTURE.md](ARCHITECTURE.md) · [RESEARCH.md](RESEARCH.md) · gate criteria in [GATES.md](GATES.md).

## License

[MIT](LICENSE)
