Metadata-Version: 2.4
Name: whsearch
Version: 0.3.0
Summary: Modular AI-native research search engine
License: GPL-3.0-or-later
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Internet :: WWW/HTTP :: Indexing/Search
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: beautifulsoup4<5,>=4.12
Requires-Dist: httpx<1,>=0.27
Requires-Dist: trafilatura<3,>=2
Provides-Extra: test
Requires-Dist: pytest<9,>=8; extra == "test"
Provides-Extra: mcp
Requires-Dist: mcp<2,>=1; extra == "mcp"
Provides-Extra: js
Requires-Dist: playwright<2,>=1.30; extra == "js"
Provides-Extra: pdf
Requires-Dist: pypdf<6,>=4; extra == "pdf"
Dynamic: license-file

# WHSearch

AI-native research search engine built as a lightweight modular monolith.

## Current phase

v0.3.0 — all phases + hardening: keyless multi-provider discovery
(DuckDuckGo lite/html + Wikipedia en/th routed + Google News + OpenAlex +
arXiv + Crossref + YouTube + optional self-hosted SearXNG), robots-gated
reader (HTML/PDF/plain/RSS, JS-shell L3 fallback), weighted RRF fusion,
BM25 passage retrieval with section boosts, adaptive multi-round planning
with evidence-gap follow-ups, Thai-aware claims/contradictions with
PSL domains, claim-level verification, MCP agent (`search`, `read_page`,
`search_and_read`, `research`), bounded TTL SQLite/FTS5 index, and explicit
research budgets.

## Design rules

- Evidence first: retrieve sources and passages before generating research conclusions.
- Online first: use external discovery initially; keep persistent storage bounded.
- Respect robots.txt, access policies, and conservative per-domain rate limits.
- Domain models and protocols must not depend on HTTP clients, providers, MCP, or extractors.
- MCP is an adapter layer; research/search logic stays in application modules.
- Optional integrations must not be required for importing the core domain.
- Resource limits are explicit so the system remains usable on low-memory machines.
- Keyless only: no paid API keys required for any default path.

## Architecture

```text
domain (models/protocols, no httpx/bs4/mcp)
  ^            ^            ^
search/reader/retrieval/research/evidence (pure app logic)
  ^            ^            ^
agent/service (plan → search → read → verify, budgeted)
  ^
application (composition root: providers, reader, index, agent)
  ^
mcp/server + tools (thin adapters) / infrastructure/http
```

## Query DSL

`site:`, `-exclude`, `"exact phrase"`, `filetype:`, `OR` (AND/NOT → spaces).
`site:` becomes a domain allowlist; `-`/`"phrase"` filter locally so flaky
HTML providers still work. `recency_days` boosts fresh dated results
(undated kept) and the agent drops stale dated pages post-read.

Thai: `th/en/mixed` detection routes to one Wikipedia edition, sets
`kl/hl/gl/ceid` + `Accept-Language`, and enables Thai-aware BM25/claims.

## Planned phases

1. ~~Foundation: contracts, configuration, logging, testing, architecture checks.~~ Done.
2. ~~Web discovery: provider abstraction and DuckDuckGo discovery.~~ Done (+Wikipedia).
3. ~~Web reader: robots policy, fetching, extraction, metadata, passages.~~ Done.
4. ~~Retrieval: passage ranking and deduplication.~~ Done.
5. ~~Research: adaptive query planning and stopping conditions.~~ Done.
6. ~~Evidence: claims, source independence, contradictions, verification.~~ Done.
7. ~~Search agent: orchestration through the MCP tools.~~ Done.
8. ~~Local index: SQLite/FTS5 bounded cache and reusable evidence.~~ Done.
9. ~~Autonomous research budgets and larger-scale discovery.~~ Done (budgets + fan-out).
10. v0.3 hardening: weighted fusion, follow-up planning, Thai evidence, PDF/RSS,
    TTL caches, PSL domains, request IDs + JSON logs. Done.

## Development

Use the repository virtual environment when available:

```bash
.venv/bin/python -m pytest
.venv/bin/python -m compileall -q src tests
.venv/bin/ruff check src tests
```

The quality gate must pass before moving to the next phase.

## Install as an MCP server

Requires Python >= 3.12. After the `whsearch` package is published to PyPI,
no manual install is needed — `uvx` fetches and runs it on first use:

```json
{
  "mcpServers": {
    "whsearch": { "command": "uvx", "args": ["--from", "whsearch[mcp]", "whsearch"] }
  }
}
```

Alternatives:

```bash
uv tool install "whsearch[mcp]" && whsearch   # persistent install via uv
pipx install "whsearch[mcp]" && whsearch      # persistent install via pipx
pip install -e ".[mcp]" && whsearch           # from source
```

`WHSEARCH_INDEX_PATH` enables the persistent local index.

## Configuration (env)

| Var | Default | Meaning |
|-----|---------|---------|
| `WHSEARCH_USER_AGENT` | `WHSearch/0.2` | Outbound UA |
| `WHSEARCH_REQUEST_TIMEOUT` | `15` | HTTP timeout (s) |
| `WHSEARCH_MAX_RESPONSE_BYTES` | `5000000` | Reader cap |
| `WHSEARCH_MAX_SEARCH_RESULTS` | `30` | Fusion cap |
| `WHSEARCH_MAX_PAGES` | `20` | Agent page budget default |
| `WHSEARCH_DOMAIN_DELAY` | `1.0` | Per-domain politeness (s) |
| `WHSEARCH_CACHE_MAX_BYTES` | `500000000` | Index byte bound |
| `WHSEARCH_INDEX_PATH` | — | Enable SQLite index |
| `WHSEARCH_INDEX_MAX_ENTRIES` | `5000` | Count bound |
| `WHSEARCH_INDEX_TTL_DAYS` | — | TTL pruning (empty = off) |
| `WHSEARCH_SEARCH_CACHE_TTL` | `600` | Positive search cache (s) |
| `WHSEARCH_SEARCH_NEGATIVE_TTL` | `60` | Empty-result cache (s) |
| `WHSEARCH_READ_CONCURRENCY` | `5` | Parallel page reads |
| `WHSEARCH_TOP_K` / `WHSEARCH_MIN_SCORE` | `3` / `0.5` | Evidence thresholds |
| `WHSEARCH_MIN_SUPPORT` / `WHSEARCH_MIN_DOMAINS` | `2` / `2` | SUPPORTED bar |
| `WHSEARCH_MIN_NEW_RATIO` / `WHSEARCH_MIN_SUPPORTED` | `0.15` / `2` | Stopping |
| `WHSEARCH_MAILTO` | `whsearch@example.com` | Polite OpenAlex/Crossref UA |
| `WHSEARCH_SEARXNG_URL` | — | Opt-in self-hosted SearXNG |
| `WHSEARCH_PDF` | `1` | `0` disables PDF parsing |
| `WHSEARCH_LOG_LEVEL` / `WHSEARCH_LOG_FORMAT` | `INFO` / `text` | `json` for structured logs |

`pip install -e ".[pdf]"` enables PDF (`pypdf`).

## Optional headless-browser fallback (L3)

Pages that render only via JavaScript (JS-shell SPAs) defeat static
extraction. When the `js` extra is installed, the reader tries headless
Chromium **only** for pages where static extraction yields almost nothing:

```bash
pip install -e ".[mcp,js]" && .venv/bin/playwright install chromium
```

Behavior and limits (all free, no keys):

- L1 trafilatura → L2 embedded JSON/meta → L3 headless, first hit wins.
- L3 triggers only below 200 extracted chars; rendered text must also clear it.
- At most 2 concurrent renders, 15s each, images/fonts/media blocked.
- Missing playwright (or any render failure) degrades silently to static text.
- `WHSEARCH_BROWSER=0` disables it; `WHSEARCH_BROWSER_TIMEOUT` tunes seconds.
- `WHSEARCH_CHROMIUM_PATH=/usr/bin/chromium` reuses a system browser instead of
  downloading one (`playwright install chromium`).

## Video vertical (YouTube, keyless)

Video-intent queries (youtube/video/คลิป, or `site:youtube.com`) also fan out
to a first-party YouTube search provider (`ytInitialData` parsing — no key,
no Invidious/Piped instances). Watch URLs read back as documents built from
the video's own metadata + description, with `MM:SS` chapter lines as passage
sections, so a multimodal caller knows *where* in the video to look:

- Direct captions are intentionally not fetched: YouTube's timedtext now
  requires proof-of-origin tokens and rate-limits keyless access (429).
- `whsearch://stats` reports `reader.browser_installed/enabled/timeout`.

## License

GPL-3.0-or-later, see `LICENSE`. Copyright (C) 2026 WHSearch contributors.
Per-file copyright holder names were intentionally left generic; update them
to your name before publishing if you are the sole author.
