Metadata-Version: 2.4
Name: py-epg
Version: 0.6.0
Summary: py_epg is an easy to use, modular, multi-process EPG grabber written in Python.
License: MIT
License-File: LICENSE.txt
Keywords: xmltv,epg
Author: Szabolcs Fruhwald
Author-email: mail@szab100.com
Requires-Python: >=3.9
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: 3.15
Requires-Dist: PyYAML (>=6.0.2,<7.0.0)
Requires-Dist: beautifulsoup4 (>=4.12.0,<5.0.0)
Requires-Dist: fake-useragent (>=2.0.0,<3.0.0)
Requires-Dist: lxml
Requires-Dist: py-xmltv (>=1.0.8,<2.0.0)
Requires-Dist: pysocks (>=1.7.1,<2.0.0)
Requires-Dist: python-dateutil (>=2.9.0,<3.0.0)
Requires-Dist: requests (>=2.32.0,<3.0.0)
Requires-Dist: roman (>=5.0,<6.0)
Requires-Dist: tqdm (>=4.67.0,<5.0.0)
Requires-Dist: xsdata (>=21.9,<22.0)
Project-URL: Homepage, https://github.com/szab100/py-epg
Project-URL: Repository, https://github.com/szab100/py-epg
Description-Content-Type: text/markdown

# py-epg

**py-epg** is an easy to use, modular, multi-process EPG grabber written in Python.

* 📺 Scrapes various TV Program websites and saves programs in XMLTV format.
* 🧩 Simply extend [EpgScraper](https://github.com/szab100/py_epg/blob/main/py_epg/common/epg_scraper.py) to grab EPG from your favorite TV site (requires basic Python skills).
* 🤖 The framework provides the rest:
    * [Beautiful Soup](https://www.crummy.com/software/BeautifulSoup/bs4/doc) - easily search & extract data from html elements 
    * multi-processing
    * config management
    * logging
    * build & write XMLTV (with auto-generated fields, eg 'stop')
    * single proxy or **rotating proxy pool** support (with remote proxy-list
      fetching, per-request rotation, timeouts, retries and a circuit
      breaker that auto-disables failing proxies)
    * **persistent caching** (SQLite) of slow-changing data (channel logos,
      program details) with configurable TTLs
    * auto http/s retries
    * random fake user_agents
* 🚀 Save time by fetching channels in parallel (caution: use proxy server(s) to avoid getting blacklisted)!
* 🧑🏻‍💻 Your contributions are welcome! Feel free to create a PR with your tv-site scraper and/or framework improvements.

<p align="center">
  <img src="https://raw.githubusercontent.com/szab100/py_epg/main/py_epg.gif">
</p>

## Usage

1. Install package:
    ```sh
    $ pip3 install py_epg
    ```
2. Create configuration: py_epg.xml
    - Add all your channels (see [sample config](https://github.com/szab100/py-epg/blob/main/py_epg.xml)).
    - Make sure there is a corresponding site scraper implementation in [py_epg/scrapers](https://github.com/szab100/py-epg/tree/main/py_epg/scrapers) for each channels ('site' attribute).
    - Optionally configure a `<cache>` (persistent caching with TTLs) and a
      `<proxy-list>` (rotating proxy pool with circuit breaker) - see the
      comments in the sample config for all supported attributes.
3. Run:
    ```sh
    $ python3 -m py_epg -c </path/to/your/py_epg.xml> -p
    ```

    ..or see all supported flags:
    ```sh
    $ python3 -m py_epg -h
    usage: py_epg [-h] [-p [PROGRESS_BAR]] [-q [QUIET]] -c CONFIG
    ...
    ```

## Development

Your contributions are welcome! Setup your dev environment as described below. [VSCode](https://code.visualstudio.com/) is a great free IDE for python projects. Once you are ready with your cool tv site scraper or framework feature, feel free to open a Pull Request here.

1. Install poetry: 
    ```sh
    curl -sSL https://raw.githubusercontent.com/python-poetry/poetry/master/get-poetry.py | python -
    ```

2. Clone repository & install dependencies:
      ```sh
      git clone https://github.com/szab100/py-epg.git && cd py-epg
      ```

3. Configure py_epg.xml
    - Add all your channels (see the sample config xml). Make sure you have a scraper implementation in py_epg/scrapers/ for each channels ('site' attribute).

4. Run:
      ```sh
      poetry install
      poetry run epg -c py_epg.xml
      ```

## License

Copyright 2021. Released under the MIT license.

