Metadata-Version: 2.4
Name: TicketMetric_tool
Version: 0.1.2
Summary: Shared internal helpers: config, logging, Mongo, retries, migrations, jobs, notifications.
License-Expression: LicenseRef-Proprietary
Project-URL: Source, https://github.com/aryansingh3/common_internal_tool
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Intended Audience :: Developers
Classifier: Topic :: Software Development :: Libraries
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: pymongo>=4.6.0
Requires-Dist: python-dotenv>=1.0.0
Requires-Dist: requests>=2.31.0
Provides-Extra: flask
Requires-Dist: Flask>=2.3.0; extra == "flask"
Provides-Extra: logging
Requires-Dist: pretty-pie-log>=0.1.0; extra == "logging"
Provides-Extra: dev
Requires-Dist: pytest>=7.4; extra == "dev"
Requires-Dist: mongomock>=4.1; extra == "dev"

# TicketMetric_tool

Shared internal helpers for the API, the scrapers, the crons and one-off migrations.

> **Two names, on purpose.** The pip distribution is
> `TicketMetric_tool`; the import is the short `tmcommon`.
> Same convention as `pip install python-dateutil` / `import dateutil`, so call
> sites stay readable instead of carrying a 36 character prefix.
>
> ```bash
> pip install TicketMetric_tool
> ```
> ```python
> from tmcommon.jobs import Job
> ```
>
> To make them identical instead, rename `src/tmcommon/` and update the two
> `[project.scripts]` paths in `pyproject.toml`.

Design rule: **the library ships mechanism, the app keeps policy and schema.** Connection
handling, batching, retries and migration plumbing live here. Collection names, index
definitions and business queries stay in the app.

It is importable with no Flask installed, no Slack configured and no database reachable,
because migrations and scrapers need it too.

## Install

```bash
# from the repo
pip install -e /path/to/common_internal_tool            # core
pip install -e ".[flask,logging]"                       # API services

# straight from GitHub
pip install "git+https://github.com/aryansingh3/common_internal_tool.git@main"
```

Pin a release once there is one, rather than tracking `main`:

```bash
pip install "git+https://github.com/aryansingh3/common_internal_tool.git@v0.1.0"
```

Build a wheel to hand around or host yourself:

```bash
python -m build --wheel        # dist/ticketmetric_tool-0.1.2-py3-none-any.whl
```

## Getting changes into the consuming repos

pip does not watch GitHub. Installing takes a snapshot, so a push to `main` changes
nothing in an environment that already installed the package. Pick the workflow that
matches where you are:

**Developing the library and an app together.** Clone once, install editable, and every
`git pull` is live with no reinstall. This is the only setup where changes really are
automatic.

```bash
git clone https://github.com/aryansingh3/common_internal_tool.git
pip install -e ../common_internal_tool
```

**Pulling the latest `main` into an environment.** pip caches aggressively, so an
`--upgrade` alone often appears to do nothing when the version string has not moved:

```bash
pip install --upgrade --force-reinstall --no-cache-dir \
  "git+https://github.com/aryansingh3/common_internal_tool.git@main"
```

**Deployments.** Pin a tag, never `main`. A deploy that resolves `main` is not
reproducible, and the same Dockerfile will produce different images on different days.

```
# requirements.txt
TicketMetric_tool @ git+https://github.com/aryansingh3/common_internal_tool.git@v0.1.2
```

Bumping that pin is then a reviewable one-line diff, which is what you want for a library
several services depend on.

### Cutting a release

```bash
# bump version in pyproject.toml, commit, then
git tag v0.2.0 && git push origin v0.2.0
```

The release workflow verifies the tag matches `pyproject.toml`, runs the suite, builds the
wheel and sdist, and attaches them to a GitHub Release.

### The private-repo gotcha

This repo is private, so anything without your local git credentials, CI, Docker builds,
production hosts, cannot clone it. A `pip install git+https://...` there fails with an
authentication error. The usual fixes:

- a deploy key on this repo, with the private key as a secret in the consumer, using the
  `git+ssh://git@github.com/...` form
- a fine-grained PAT with read access, injected as
  `git+https://${TOKEN}@github.com/...`
- or build the wheel in CI and push it to a private index

For Docker, mount the credential as a build secret rather than baking a token into a layer.

## Python support

**3.9 through 3.13.** 3.9 is the floor because the Flask API runs on it, so nothing here
uses 3.10+ syntax: no PEP 604 `X | Y` unions, no `match` statements. CI compiles every
file on each version to enforce that, including modules the tests do not import.

## Tests

```bash
python3 run_tests.py
```

No pytest needed, so the suite runs on an interpreter carrying only the core
dependencies. The files are valid pytest modules too. A test file whose optional
dependency is missing prints `SKIP` and passes, so no Flask is not a failure.

Verified locally on 3.9.6 and 3.11.15; 3.10, 3.12 and 3.13 are covered by the CI matrix.

`tests/fake_mongo.py` is a small in-memory `Collection` covering only the operations these
modules use, so lock and batch logic is testable without mongomock and without pointing a
test at a real cluster.

## Environment

| variable | required | purpose |
|---|---|---|
| `SERVICE_NAME` | recommended | identifies the service in logs and alerts |
| `APP_ENV` | no | `dev`/`local`/`test`/`staging` suppress Slack; anything else is production |
| `LOG_DIR` | no | absolute log directory, defaults to `<cwd>/logs` |
| `LOG_LEVEL` | no | defaults to `DEBUG` |
| `MONGODB_URI` | when using mongo | default cluster |
| `MONGODB_URI_<ALIAS>` | no | extra clusters, e.g. `MONGODB_URI_STAGING` |
| `SLACK_WEBHOOK_URL` | when notifying | absent means notifications are skipped, not an error |

## Modules

| module | what it gives you |
|---|---|
| `config` | typed env access that raises at point of use, not on import |
| `logging` | `get_logger(name)`, cached, absolute log dir, falls back to stdlib when `pretty_pie_log` is absent |
| `context` | correlation id and extra fields via ContextVar, so a cron traceback is traceable |
| `exceptions` | `BaseAPIException` and the HTTP subclasses, dependency-free |
| `serialization` | one JSON encoder for ObjectId and datetime, plus `to_jsonable` |
| `retry` | `@retry` with exponential backoff and jitter |
| `timing` | `timed()`, `@timeit`, and `Stopwatch` for multi-phase jobs |
| `mongo.connection` | lazy multi-cluster registry, `DatabaseRouter` for the live/historical split |
| `mongo.batch` | chunked `insert_many` / `bulk_write` that separates duplicate keys from real errors |
| `mongo.iterate` | `iter_by_id` keyset paging, no cursor timeouts, resumable |
| `mongo.checkpoint` | watermarks so incremental jobs resume instead of guessing a `--since` |
| `migrations` | versioned migrations with an applied record, a lock and dry-run |
| `jobs` | scheduled work: expiring locks, run history, heartbeats, overdue detection |
| `notify.slack` | error alerts, framework-agnostic, context supplied by the app |
| `flask_ext` | `init_app(app)` for correlation ids and error handlers |

## Jobs and crons

A `Job` gets a lock, a run row, a status document with heartbeat, a correlation id, a
checkpoint, failure alerting and timing. Records use the field names already in
`vc_worker_runs` and `vc_worker_status`, so existing queries keep working.

```python
from tmcommon.jobs import Job

class SectionRosterJob(Job):
    name = 'tm_section_roster'
    description = 'Fold new TM snapshots into the section roster'
    lock_ttl_seconds = 1800

    def run(self, ctx):
        since = ctx.checkpoint.get()
        processed = 0
        for event_id in events_since(since):
            update_roster(event_id)
            processed += 1
            if not ctx.heartbeat(processed=processed):
                break          # lease lost, another process took over
        ctx.checkpoint.advance(newest_seen)
        return {'events': processed}
```

Run it in-process, replacing the hand-rolled worker threads:

```python
from tmcommon.jobs import Scheduler

scheduler = Scheduler(db)
scheduler.add(SectionRosterJob(), interval_seconds=86400)
scheduler.add(CleanupJob(), interval_seconds=86400, initial_delay_seconds=300)
scheduler.start()
```

Or as a one-shot under system cron or a Kubernetes CronJob, with the same recording:

```bash
tm-job run     --package jobs --db tickets --name tm_section_roster
tm-job status  --db tickets
tm-job overdue --db tickets --expect tm_section_roster=86400,cleanup_worker=86400
```

`overdue` exits non-zero when a job is late, so a monitor can alert on it. That is the
failure currently invisible: `cron.py` is a bare `while True` loop, and if that process
dies every schedule inside it stops with no signal.

**Why the lock matters here.** `app.py` starts the workers at module scope, in the `else`
branch gunicorn takes, so every web worker process starts its own copy. Four gunicorn
workers run four cleanup loops concurrently today. Under `Scheduler` all four still start
and exactly one does the work per tick; the rest record `skipped`.

## Flask wiring

```python
from tmcommon.flask_ext import init_app, flask_context_provider
from tmcommon.notify import register_context_provider

init_app(app)
register_context_provider(lambda: {**flask_context_provider(),
                                   'User': get_request_user().email})
```

`BaseAPIException` becomes its `to_dict()` response and is logged at warning. Anything else
is treated as a bug: logged with a traceback, sent to Slack, returned as a 500. A 404 or 405
is neither.

## Migrations

```python
# migrations/m20261001_section_roster.py
from tmcommon.migrations import Migration

class SectionRoster(Migration):
    version = '20261001_section_roster'
    description = 'Build tm_section_roster from ticketmaster_detail_data'

    def up(self, db, dry_run=False):
        count = db.ticketmaster_detail_data.estimated_document_count()
        if dry_run:
            return {'would_scan': count}
        ...
        return {'events': 412, 'sections': 9304}
```

```bash
tm-migrate status --package migrations --db tickets
tm-migrate up     --package migrations --db tickets            # dry run
tm-migrate up     --package migrations --db tickets --apply
```

Writes need `--apply`. Deliberate, given these run against collections with millions of rows.

## Deliberately not here

- **Collection accessors and index definitions.** Your `DatabaseConnection._create_indexes()`
  is app schema. Declare it in the app and call it at startup.
- **`DatabaseManager` business queries.** `get_user_from_email`, `get_analytics` and the
  venue lookups are domain logic, not shared mechanism.
- **Domain models** such as `APIUser` and `Membership`.

## Roadmap

Not built yet, in rough priority order:

1. `http` client: session with retry, proxy rotation and UA rotation. Both repos already carry
   `proxies.json` and `user_agent.json` plus their own rotation code.
2. `pipeline` base: extract/transform/load with stats and failure reporting, to replace the
   per-spider boilerplate.
3. `ratelimit`: token bucket for upstream politeness.
4. `dates`: `est_to_utc`, `convert_utc_to_timezone`, currently duplicated in both repos.
5. `mongo.upsert`: find-or-create keyed on an alternate identity, the gap behind the duplicate
   events bug.
6. `schema`: declarative model base with `to_dict` / `from_dict` and validation.
7. Test helpers: `mongomock` fixtures and a fake clock.
