Metadata-Version: 2.5
Name: pipscout
Version: 0.1.0
Summary: Tell it what you need in plain English. Get the best PyPI package for the job.
Project-URL: Homepage, https://github.com/Meet2147/pythonLibraries/tree/main/pipscout
Project-URL: Repository, https://github.com/Meet2147/pythonLibraries
Project-URL: Issues, https://github.com/Meet2147/pythonLibraries/issues
Author-email: Meet2147 <meetjethwa3@gmail.com>
License-Expression: MIT
License-File: LICENSE
Keywords: cli,discovery,packages,pip,pypi,recommendation,search
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: System :: Software Distribution
Classifier: Topic :: Utilities
Requires-Python: >=3.9
Requires-Dist: rich>=13
Provides-Extra: dev
Requires-Dist: pytest>=7; extra == 'dev'
Description-Content-Type: text/markdown

<p align="center">
  <img src="https://raw.githubusercontent.com/Meet2147/pythonLibraries/main/pipscout/assets/banner.svg" alt="pipscout — say what you need, get the best PyPI package" width="760">
</p>

<p align="center">
  <a href="https://pypi.org/project/pipscout/"><img src="https://img.shields.io/pypi/v/pipscout.svg?color=a78bfa&label=pypi" alt="PyPI"></a>
  <img src="https://img.shields.io/badge/python-3.9%2B-22d3ee" alt="Python 3.9+">
  <a href="https://github.com/Meet2147/pythonLibraries/actions/workflows/pipscout-ci.yml"><img src="https://github.com/Meet2147/pythonLibraries/actions/workflows/pipscout-ci.yml/badge.svg" alt="CI"></a>
  <img src="https://img.shields.io/badge/deps-just%20rich-f472b6" alt="Dependencies: just rich">
  <img src="https://img.shields.io/badge/license-MIT-4ade80" alt="MIT license">
</p>

<p align="center">
  <b>Stop googling “best python library for X”.</b><br>
  Tell <code>pipscout</code> what you're building. It combs through PyPI and hands you the package everyone actually uses — with receipts.
</p>

---

```bash
pip install pipscout
pipscout "pdf extraction"
```

<p align="center">
  <img src="https://raw.githubusercontent.com/Meet2147/pythonLibraries/main/pipscout/assets/demo-pdf.svg" alt="pipscout recommending pymupdf, pypdf, pdfminer.six and pdfplumber for 'pdf extraction'" width="100%">
</p>

That's it. 🔭 One command gets you the **best pick**, the **runners-up**, their **monthly downloads**, how **fresh** each release is, and the exact `pip install` line.

## ✨ Why you'll like it

| | |
|---|---|
| 🗣️ **Plain English in** | `"pdf extraction"`, `"i want to parse yaml files"`, `"excel spreadsheets"` — it strips the filler, knows `extraction` ≈ `extract` ≈ `parse`, and gets to work. |
| 🏆 **The favourite out** | Ranks by what matters: *is it on-topic*, *does the world actually use it* (30-day downloads), *is it still maintained*. Abandoned and “Inactive” projects sink. |
| 🧾 **Receipts, not vibes** | Every pick shows *why*: which words matched, downloads per month, last release, relevance & freshness meters. |
| ⚡ **Fast** | ~2 s on a cold cache, near-instant after — metadata is cached for a week and fetched 16-wide in parallel. |
| 🪶 **Featherweight** | One dependency (`rich`, for the drip). Ships a 170 KB catalog of PyPI's top 5,000 projects so it's smart from the very first run. |
| 🤖 **Script-friendly** | `--json` output and a one-function Python API. |

## 🚀 Take it for a spin

```bash
pipscout "web scraping"                  # 🥇🥈🥉 top 5
pipscout "data validation" -n 10         # show me more
pipscout "http client" --json | jq       # pipe it anywhere
pipscout "pdf extraction" --deep         # also scan the names of all ~700k projects on PyPI
pipscout index                           # optional: pre-index the top 5,000 so READMEs are searchable too
```

<p align="center">
  <img src="https://raw.githubusercontent.com/Meet2147/pythonLibraries/main/pipscout/assets/demo-scraping.svg" alt="pipscout recommending beautifulsoup4 for 'web scraping'" width="100%">
</p>

> 👀 `beautifulsoup4` won “web scraping” even though neither word is in its name — pipscout reads summaries and keywords (and, after `pipscout index`, READMEs), not just names.

### Real results, fresh install, no index

| You type | pipscout says |
|---|---|
| `pdf extraction` | **pymupdf** → pypdf → pdfminer.six → pdfplumber |
| `web scraping` | **beautifulsoup4** → htmldate → trafilatura → firecrawl-py → Scrapy |
| `parse yaml files` | **PyYAML** → ruamel.yaml |
| `data validation` | **pydantic** → jsonschema → email-validator |
| `http client` | **aiohttp** → httpx → httpcore → urllib3 |
| `excel spreadsheets` | **xlrd** → openpyxl → xlsxwriter → gspread |
| `plotting charts` | **matplotlib** → pyecharts |

## 🐍 From Python

```python
from pipscout import recommend

for rec in recommend("pdf extraction", limit=2):
    print(f"{rec.name:<8} {rec.score:.2f}  {rec.downloads or 0:>12,} dl/mo")
    print("   ", rec.install, "|", "; ".join(rec.reasons))
```

```text
pymupdf  0.46   101,779,570 dl/mo
    pip install pymupdf | summary matches 'pdf'; summary matches 'extract'; 101.8M downloads in the last 30 days
pypdf    0.30   151,277,866 dl/mo
    pip install pypdf | summary matches 'pdf'; description matches 'extract'; 151.3M downloads in the last 30 days
```

Each `Recommendation` has `name`, `score`, `relevance`, `popularity`, `health`, `downloads`, `summary`, `version`, `last_release`, `homepage`, `install`, `reasons`, `warnings` — plus `.to_dict()` for JSON.

## 🧠 How the magic works

PyPI's website search has no public API (and sits behind a bot wall), so pipscout brings its own brain:

```
  "i want pdf extraction"
          │
   1. 🗣️  understand ── drop filler · stem (extraction → extract) · add synonyms (extract ≈ parse ≈ read)
          │
   2. 🔎  shortlist ─── names of the ~15k most-downloaded projects
          │             + summaries & keywords from the bundled top-5k catalog
          │             + READMEs from your local index · + every name on PyPI with --deep
          │
   3. 📡  inspect ───── live metadata for the best 60 from PyPI's JSON API (parallel, cached 7 days)
          │
   4. 🏆  rank ──────── score = relevance² × popularity × √freshness
```

- **relevance** (0–1) — how many of your concepts a package covers, each weighted by how *rare* the word is on PyPI (so `pdf` outweighs `data`). A hit in the name or summary counts more than one buried in the README; a synonym counts a bit less than the real word. Anything under 0.5 is dropped.
- **popularity** — `(monthly downloads ÷ 1B)^¼`. A package used 10× more can beat a slightly more on-topic one… but never an off-topic one.
- **freshness** — 1.0 if released in the last year, sliding to 0.2 at five years, with an extra haircut for *Inactive* or *Alpha* status.

Download counts come from the excellent [top-pypi-packages](https://github.com/hugovk/top-pypi-packages) dataset (PyPI's public BigQuery stats, refreshed monthly).

## 🎛️ CLI reference

| Command / flag | What it does |
|---|---|
| `pipscout "<need>"` | Recommend packages (shorthand for `pipscout search "<need>"`) |
| `-n, --limit N` | How many results (default 5) |
| `--json` | Machine-readable output |
| `--deep` | Also match names across every project on PyPI (fetches the ~45 MB name index, cached for a day) |
| `--min-relevance X` | Loosen or tighten the on-topic bar (0–1, default 0.5) |
| `--max-fetch N` | Candidates to inspect in detail (default 60) |
| `--refresh` | Ignore the cache and re-fetch |
| `pipscout index [--top N]` | Pre-fetch metadata for the top N packages (default 5,000, ~3 min) |
| `pipscout clear-cache` | Wipe the cache (`~/.cache/pipscout`, or `$PIPSCOUT_CACHE_DIR`) |

## 🤔 FAQ

**Is it just sorting by downloads?** Nope. Downloads only count once a package is on-topic — relevance is squared, so an off-topic giant (looking at you, `requests`) never wins a PDF query.

**Why not just use PyPI search?** It sorts by text match, not by what people actually use, has no API, and won't tell you a project died in 2017.

**Does it phone home?** Only to `pypi.org` and `raw.githubusercontent.com` (for the download stats). No telemetry, no accounts, no API keys.

**It missed my favourite package!** Try `--deep`, rephrase with the words the package would use to describe itself, or run `pipscout index` once so READMEs get searched too. PRs to the synonym list in [`text.py`](https://github.com/Meet2147/pythonLibraries/blob/main/pipscout/src/pipscout/text.py) are very welcome.

## 🛠️ Hacking on it

```bash
git clone https://github.com/Meet2147/pythonLibraries && cd pythonLibraries/pipscout
pip install -e ".[dev]"
pytest                                                   # fully offline — PyPI is faked
python scripts/render_demo.py                            # re-shoot the terminal screenshots
pipscout index && python scripts/build_bundled_data.py   # refresh the bundled catalog
```

## 📜 License

MIT © [Meet2147](https://github.com/Meet2147) — now go build something. 🔭
