Metadata-Version: 2.4
Name: zytools-fs
Version: 0.0.16
Summary: Personal Python utilities for FTP, video, and audio downloads
Author: zytools
License-Expression: MIT
Keywords: zytools,ftp,video,download
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE.txt
Requires-Dist: cryptography
Requires-Dist: lmdb
Requires-Dist: loguru
Requires-Dist: lxml
Requires-Dist: paramiko
Requires-Dist: requests
Requires-Dist: tqdm
Requires-Dist: trafilatura
Dynamic: license-file

﻿# zytools-fs

`zytools-fs` is a small collection of Python utilities for personal crawling,
FTP transfer, task heartbeats, persistent URL de-duplication, Bilibili video
downloads, and Maoer FM audio downloads.

The package is intentionally lightweight: every helper can be imported and used
directly in scripts without a framework.

## Features

- FTP and SFTP recursive download/upload with retry, progress, skip-by-size, and concurrent transfer support.
- Bilibili video downloader based on `requests`.
- Maoer FM audio downloader with DRM decryption and WAV output.
- LMDB-backed URL filter for large crawl de-duplication.
- Article page detector and simple same-domain crawler.
- Task heartbeat helper for reporting script status to a `/tasks` endpoint.
- `zytools` command for a quick installation check.

## Installation

```bash
pip install zytools-fs
```

Python 3.9 or newer is required.

For Bilibili DASH merging and Maoer audio conversion, install `ffmpeg` and make
sure it is available in `PATH`. Without it, Maoer downloads cannot produce the
final audio file.

## Quick Check

```bash
zytools
```

Expected output:

```text
zytools installed successfully.
```

## Release Notes

Version `0.0.16` adds Maoer FM audio downloads and fixes several correctness
and security issues across the package:

- Preserve and concatenate every Bilibili `durl` segment instead of downloading
  only the first segment.
- Parse Maoer HLS initialization and media byte ranges, reject incomplete
  encrypted segments, and avoid logging DRM keys.
- Restore persistent URL de-duplication for newly created empty filters.
- Verify HTTPS certificates during article crawling.
- Validate FTP and SFTP download sizes before replacing local files.
- Handle non-JSON task-service responses without unexpectedly raising when
  `raise_error=False`.

Complete release details are available in `CHANGELOG.md` in the source
distribution.

## FTP and SFTP Transfer

`FTPClient` and `SFTPClient` share the same high-level API for recursive
upload/download. Both clients support per-file progress logs, retries,
skip-by-size, concurrent directory transfers, and final transfer summaries.

### FTP

```python
from zytools.utils import FTPClient

with FTPClient(
    host="127.0.0.1",
    user="user",
    password="password",
    port=21,
    passive=True,
    workers=4,
) as ftp:
    ftp.download("/remote/path", "./downloads")
    ftp.upload("./reports", "/remote/reports")
```

### SFTP

```python
from zytools.utils import SFTPClient

with SFTPClient(
    host="127.0.0.1",
    user="user",
    password="password",
    port=22,
    # key_filename="~/.ssh/id_rsa",
    # allow_unknown_host=True,  # only use this for trusted private hosts
    workers=4,
) as sftp:
    download_result = sftp.download("/remote/path", "./downloads")
    upload_result = sftp.upload("./reports", "/remote/reports")
```

### Transfer Behavior

- Direct file paths transfer one file; directory paths are traversed recursively.
- Existing target files with the same size are skipped and counted as success.
- Uploads and downloads use fixed-position progress bars when `show_progress=True`.
  Single-thread transfers use one dynamic line; concurrent transfers reuse up to
  `workers` progress lines instead of printing endless progress logs.
- Each `download()` or `upload()` call logs and returns `total`, `success`, and `error`.
- `workers>1` enables concurrent file transfers for directories; each worker uses
  its own connection.

Example summary and return value:

```text
upload summary: total=10 success=9 error=1
download summary: total=10 success=10 error=0
```

```python
{"total": 10, "success": 9, "error": 1}
```

Useful options:

- `port`: FTP defaults to `21`; SFTP defaults to `22`.
- `encoding`: FTP filename encoding, default `utf-8`.
- `passive`: FTP passive mode, default `True` (FTP only).
- `key_filename`: SSH private key path for SFTP authentication; `~` is expanded.
- `allow_unknown_host`: allow unknown SFTP host keys, default `False`.
- `download_retries` and `upload_retries`: retry count.
- `retry_wait_seconds`: wait time between retries.
- `workers`: concurrent workers for directory downloads and uploads, default `1`.
- `show_progress`: show fixed-position progress bars and completion logs.

## Bilibili Video Download

```python
from zytools.video import download_bili_video

ok = download_bili_video(
    "https://www.bilibili.com/video/BVxxxx",
    output_dir="./downloads",
    quality="max",
    page="all",
    filename="Bilibili_{BV}_{Date}_{Page}_{PartTitle}",
    cookie={
        "SESSDATA": "your_sessdata",
        "bili_jct": "your_bili_jct",
    },
    proxies={"https": "http://127.0.0.1:7890"},
)

print(ok)
```

Parameters:

- `quality`: `"max"` for the highest available stream, `"min"` for the lowest.
  The actual quality depends on Bilibili account permissions and returned DASH streams.
- `page`: `"all"`, a single page such as `"1"`, or a range/list such as
  `"1,3-5"`.
- `cookie`: optional Bilibili cookies for videos that require login.
- `proxies`: optional proxies for Bilibili page/API requests; media stream downloads do not use it.
- `force`: when `False`, existing final video files are skipped; when `True`, they are overwritten.

Filename template fields:

- `{Title}`: video title.
- `{BV}`: BV id.
- `{Date}`: publish date in `YYYYMMDD` format.
- `{Page}`: page number.
- `{Part}`: same as page number.
- `{Duration}`: duration in seconds.
- `{PartTitle}`: page title.

Only download content that you own or are allowed to download.

## Maoer FM Audio Download

```python
from zytools.video import download_maoer_video

output_file = download_maoer_video(
    sound_id=13073155,
    filepath="./downloads",
    proxies={
        "http": "http://127.0.0.1:7890",
        "https": "http://127.0.0.1:7890",
    },
    output_format="wav",
)
```

`filepath` is the output directory and defaults to `./downloads`. Supported
formats are `wav`, `m4a`, `mp3`, and `flac`; the default is `wav`. Proxies are
used only for page, playlist, and DRM API requests. Audio segment downloads
bypass both supplied and environment proxies. The function returns the absolute
path of the completed audio file.

Only download content that you own or are allowed to download.

## URL Filter

`UrlFilter` stores compact MD5 fingerprints in LMDB. Unlike a Bloom filter, it
does not intentionally produce false positives.

```python
from zytools.utils import UrlFilter

with UrlFilter(file_path="url_seen.lmdb") as url_filter:
    url = "https://example.com/video?id=1"

    if url_filter.add(url):
        print("new url")
    else:
        print("seen before")

    print(len(url_filter))
```

Batch import and export:

```python
from zytools.utils import UrlFilter

with UrlFilter("url_seen.lmdb") as url_filter:
    added = url_filter.add_many(
        [
            "https://example.com/a",
            "https://example.com/b",
        ]
    )
    url_filter.to_csv("url_seen.csv")

print(f"added {added} urls")

UrlFilter.to_lmdb("url_seen.csv", file_path="url_seen_copy.lmdb")
```

## Article Detection

Use `check_response` to request one URL and classify it as an article, other
HTML page, binary resource, or fetch error.

```python
from zytools.artice import check_response

result = check_response("https://example.com/news/1.html")

if result["type"] == "article":
    print(result["title"])
    print(result["date"])
    print(result["text"][:300])
else:
    print(result["type"], result.get("reason"))
```

Return `type` values:

- `article`: article page with extracted `title`, `date`, `author`, and `text`.
- `other`: HTML page that does not look like an article.
- `binary`: image, PDF, JavaScript, CSS, video, archive, or other non-HTML file.
- `fetch_error`: request failed or returned a bad HTTP status.

## Simple URL Crawler

`UrlCrawler` starts from one URL, follows links breadth-first, and yields article
items. It can optionally use `UrlFilter` to avoid saving the same article URL
across runs.

```python
from zytools.artice import UrlCrawler
from zytools.utils import UrlFilter

with UrlFilter("article_urls.lmdb") as url_filter:
    crawler = UrlCrawler(
        start_url="https://example.com/",
        max_saved_urls=20,
        same_domain=True,
        max_depth=5,
        url_fp=url_filter,
    )

    for item in crawler.crawl():
        print(item["title"], item["url"])

    crawler.save_url_filter()
```

Each yielded item has:

- `title`: extracted article title.
- `creat_date`: extracted article publish date.
- `content`: extracted article text.
- `url`: final article URL.
- `get_date`: crawl batch date.

## Task Heartbeat

Use `update_task` for a single heartbeat request, or `TaskUpdater` when a script
needs repeated updates with a minimum interval.

```python
from zytools.utils import TaskUpdater, update_task

result = update_task(
    name="daily job",
    machine_id="machine-1",
    script_path="/path/to/script.py",
    server="http://127.0.0.1:8001",
)

print(result)

task = TaskUpdater(
    name="daily job",
    machine_id="machine-1",
    script_path="/path/to/script.py",
    server="http://127.0.0.1:8001",
    min_interval=60,
)

task.update()
task.update(force=True)
```

The server is expected to accept `POST /tasks` with a JSON body containing
`name`, `machine_id`, `script_path`, `enabled`, and `timeout_seconds`.

## Development

Build the package locally:

```bash
python -m build
```

Check the distribution metadata:

```bash
python -m twine check dist/*
```

Publish to PyPI:

```bash
python -m twine upload dist/*
```

## License

MIT
