Metadata-Version: 2.4
Name: pulselog
Version: 2.0.0
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: System :: Logging
Classifier: Typing :: Typed
Requires-Dist: websockets>=11.0
Requires-Dist: tomli>=2.0 ; python_full_version < '3.11'
Requires-Dist: pytest>=7.0 ; extra == 'dev'
Requires-Dist: pytest-cov>=4.0 ; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.21 ; extra == 'dev'
Provides-Extra: dev
Summary: Real-time browser dashboard for Python logging — zero config, non-blocking
Keywords: logging,dashboard,websocket,real-time,monitoring
License: MIT
Requires-Python: >=3.8
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM

# ⚡ PulseLog

**A non-blocking Python logging library with a real-time browser dashboard.**

Built for ML training, data pipelines, backend services, experiments, and long-running workloads where logging should stay out of the critical path.

[![PyPI](https://img.shields.io/badge/pip%20install-pulselog-blue)](https://pypi.org/project/pulselog/)
[![License: MIT](https://img.shields.io/badge/license-MIT-green)](#license)
[![Python](https://img.shields.io/badge/python-%E2%89%A53.8-yellow)](#requirements)

---

## ✨ Why PulseLog?

Traditional logging can become surprisingly expensive when it is called inside:

- ML training loops
- ETL/data pipelines
- inference workloads
- batch processing
- concurrent worker systems
- long-running experiments

PulseLog is designed around a simple principle:

> **Logging should not become the bottleneck of the application.**

Log records are handled asynchronously so application code does not need to wait for dashboard delivery or other downstream processing.

---

## 📦 Installation

```bash
pip install pulselog
```

Then:

```python
from pulselog import Logger

log = Logger("my-app")
log.info("application started")
```

---

## 🚀 Quick Start

```python
from pulselog import Logger

log = Logger("training")

log.info("training started", epoch=1)
log.warning("learning rate is high", learning_rate=0.1)

log.shutdown()
```

When the dashboard is enabled, PulseLog provides a browser-based view of the log stream.

Default dashboard:

```text
http://localhost:5678
```

---

## 💡 Examples

### Basic Logging

```python
from pulselog import Logger

log = Logger("my-app")

log.debug("debugging info")           # DEBUG level
log.info("something happened")        # INFO level
log.warning("something seems off")    # WARNING level
log.error("something went wrong")     # ERROR level
log.critical("system failure")        # CRITICAL level

log.shutdown()
```

### With Structured Data

```python
from pulselog import Logger

log = Logger("api-server")

log.info(
    "request handled",
    method="GET",
    path="/users/42",
    status=200,
    latency_ms=12,
)

log.warning(
    "slow request",
    path="/reports",
    latency_ms=2500,
)

log.shutdown()
```

### Error Handling

```python
from pulselog import Logger

log = Logger("data-processor")

try:
    result = process_data(raw_input)
except ValueError as e:
    log.error("invalid input format", error=str(e))
except Exception:
    log.exception("unexpected error during processing")

log.shutdown()
```

### Without Dashboard (Production)

```python
from pulselog import Logger

log = Logger(
    "production-worker",
    dashboard=False,              # Disable browser dashboard
    queue_size=50000,             # Larger queue for high throughput
    worker_interval=0.005,        # Faster processing
)

for item in work_queue:
    log.info("processing", item_id=item.id)

log.shutdown()
```

### ML Training Loop

```python
from pulselog import Logger

log = Logger("training", checkpoint_path=":memory:")

for epoch in range(10):
    train_loss = train_one_epoch()
    val_accuracy = validate()

    log.info(
        f"epoch {epoch + 1} complete",
        loss=round(train_loss, 4),
        accuracy=round(val_accuracy, 4),
    )

    log.save(
        f"epoch-{epoch + 1}",
        {"loss": train_loss, "accuracy": val_accuracy},
        status="DONE",
        progress=(epoch + 1) * 10,
    )

log.shutdown()
```

### Web Request Handler

```python
from pulselog import Logger

log = Logger("web-server")

def handle_request(request):
    log.info("request received", method=request.method, path=request.path)

    try:
        response = process(request)
        log.info("request complete", status=response.status_code)
        return response

    except AuthError:
        log.warning("authentication failed", ip=request.ip)
        raise

    except Exception:
        log.exception("request failed")
        raise

log.shutdown()
```

### Background Tasks

```python
from concurrent.futures import ThreadPoolExecutor
from pulselog import Logger

log = Logger("task-runner")

def run_task(task):
    log.info("task started", task_id=task.id)
    result = task.execute()
    log.info("task finished", task_id=task.id, duration_ms=result.duration)
    return result

with ThreadPoolExecutor(max_workers=8) as pool:
    results = list(pool.map(run_task, tasks))

log.info("all tasks complete", total=len(results))
log.shutdown()
```

### Using Tags

```python
from pulselog import Logger

log = Logger("pipeline")

log.tag("ingestion")
log.info("loading data", source="database")
log.info("loaded rows", count=15000)

log.tag("transform")
log.info("applying transforms")
log.info("transforms complete")

log.tag("export")
log.info("writing output", destination="s3://bucket/data")

log.shutdown()
```

### Context Manager for Tags

```python
from pulselog import Logger

log = Logger("pipeline")

with log.context(tag="extract"):
    log.info("connecting to source")
    data = extract()
    log.info("extraction complete", rows=len(data))

with log.context(tag="transform"):
    log.info("cleaning data")
    clean = transform(data)
    log.info("transform complete")

with log.context(tag="load"):
    log.info("writing to warehouse")
    load(clean)
    log.info("load complete")

log.shutdown()
```

### Checkpoints for Resumable Work

```python
from pulselog import Logger

log = Logger("batch-job")

# Check if we already completed this step
if log.load("step-2"):
    log.info("step-2 already done, skipping")
else:
    log.info("starting step-2")
    process_step_2()
    log.save("step-2", {"status": "complete"}, status="DONE", progress=50)

log.shutdown()
```

### Standard Library Integration

```python
import logging
from pulselog.handler import PulseHandler

# Route standard logging to PulseLog
handler = PulseHandler("my-app")
logging.getLogger().addHandler(handler)
logging.getLogger().setLevel(logging.INFO)

# Now standard logging calls go through PulseLog
logging.info("application started")
logging.warning("disk space low")
logging.error("connection failed")

try:
    risky_operation()
except Exception:
    logging.exception("operation failed")  # Includes traceback

# Clean up
handler.shutdown()
```

### Monitoring Drops

```python
from pulselog import Logger
import time

log = Logger("high-throughput", queue_size=1000)

for i in range(1_000_000):
    log.info("processing", index=i)

log.flush()

stats = log.stats()
print(f"Processed: {stats.get('records_processed', 'N/A')}")
print(f"Dropped:   {stats.get('records_dropped', 'N/A')}")

if stats.get("records_dropped", 0) > 0:
    print("Consider increasing queue_size or reducing log volume")

log.shutdown()
```

### Custom Worker Interval

```python
from pulselog import Logger

# Low latency (for real-time dashboards)
log = Logger("realtime", worker_interval=0.001)  # 1ms

# Low CPU (for background batch jobs)
log = Logger("batch", worker_interval=0.1)  # 100ms

# Default (good balance)
log = Logger("default", worker_interval=0.01)  # 10ms

log.shutdown()
```

---

## 🖥️ Real-Time Dashboard

PulseLog includes a browser-based dashboard designed for real-time visibility into application logs.

It provides:

- real-time log updates
- severity filtering
- full-text search
- structured metadata
- automatic scrolling
- session export

Typical levels:

```text
DEBUG
INFO
WARNING
ERROR
CRITICAL
```

---

## 📊 Structured Logging

PulseLog supports structured metadata without requiring you to build formatted log strings manually.

```python
log.info(
    "request completed",
    request_id="abc123",
    user_id=42,
    latency_ms=18,
    status_code=200,
)
```

Structured fields are useful for data pipelines, ML experiments, API services, batch jobs, model evaluation, debugging, and operational monitoring.

---

## 🧵 Concurrent Logging

PulseLog is designed for applications where multiple threads produce logs concurrently.

```python
from concurrent.futures import ThreadPoolExecutor
from pulselog import Logger

log = Logger("worker")

def process(i):
    log.info("processing item", item=i)

with ThreadPoolExecutor(max_workers=8) as executor:
    list(executor.map(process, range(10_000)))

log.shutdown()
```

---

## 💾 Checkpoints

PulseLog includes a checkpoint store for long-running workflows.

Useful for:

- ML training
- ETL jobs
- experiments
- batch processing
- resumable workflows

```python
log.save(
    name="epoch-5",
    data={
        "loss": 0.31,
        "accuracy": 0.94,
    },
    status="DONE",
    note="best model so far",
    progress=50,
)
```

Supported statuses:

```text
DONE
IN_PROGRESS
FAILED
SKIPPED
```

Load:

```python
result = log.load("epoch-5")
```

List:

```python
names = log.checkpoints()
```

Delete:

```python
log.delete_checkpoint("epoch-3")
```

For tests or short-lived workloads:

```python
log = Logger("test", checkpoint_path=":memory:")
```

---

## 🏷️ Tags and Context

```python
log.tag("training")

log.info("epoch started")
log.info("batch processed")
```

Or use a context manager:

```python
with log.context(tag="validation"):
    log.info("validation started", dataset="test")
    log.info("validation complete", accuracy=0.94)
```

---

## ➗ Dividers

```python
log.divider("epoch boundary")
```

---

## 🔌 Standard Library Logging Integration

PulseLog provides a bridge for Python's standard `logging` package.

```python
import logging
from pulselog.handler import PulseHandler

handler = PulseHandler("my-app")
logging.getLogger().addHandler(handler)

logging.info("application started")
```

Exception information can also be forwarded:

```python
try:
    result = model.predict(data)
except Exception:
    logging.exception("prediction failed")
```

---

## 📝 Logging Methods

```python
log.debug("debug message")
log.info("information")
log.warning("warning")
log.error("error")
log.critical("critical")
```

Structured fields:

```python
log.info(
    "training batch complete",
    epoch=10,
    batch=100,
    loss=0.21,
)
```

Exception logging:

```python
try:
    result = model.predict(data)
except Exception:
    log.exception(
        "prediction failed",
        input_shape=str(data.shape),
    )
```

---

## 🔄 Flush and Shutdown

Flush pending records:

```python
success = log.flush(timeout=2.0)
```

Shutdown cleanly:

```python
log.shutdown()
```

Explicit shutdown is recommended for long-running applications so pending records can be processed before termination.

---

## 📈 Runtime Statistics

PulseLog can expose runtime statistics through:

```python
stats = log.stats()
```

Depending on configuration, statistics can include:

- records processed
- records dropped
- queue size
- queue capacity
- queue utilisation
- checkpoint activity
- dashboard clients
- uptime

---

## ⚙️ Configuration

Configuration can be supplied through:

1. `Logger()` arguments
2. environment variables
3. `pulselog.toml`
4. built-in defaults

Example:

```python
log = Logger(
    name="my-app",
    host="localhost",
    port=5678,
    auto_open=True,
    dashboard=True,
    checkpoint_path=".pulselog/checkpoints.db",
    level="DEBUG",
    worker_interval=0.01,
)
```

### Environment variables

```bash
export PULSELOG_DASHBOARD=false
export PULSELOG_HOST=0.0.0.0
export PULSELOG_PORT=8080
export PULSELOG_AUTO_OPEN=false
export PULSELOG_CHECKPOINT_PATH=/data/checkpoints.db
export PULSELOG_LEVEL=INFO
export PULSELOG_WORKER_INTERVAL=0.01
```

### `pulselog.toml`

```toml
[pulselog]

host = "0.0.0.0"
port = 8080

auto_open = false
dashboard = true

level = "INFO"
worker_interval = 0.01
checkpoint_path = ".pulselog/checkpoints.db"
```

---

## 🏭 Production Usage

For production environments where the browser dashboard is unnecessary:

```python
from pulselog import Logger

log = Logger(
    "production",
    dashboard=False,
    checkpoint_path="/data/checkpoints.db",
)
```

For CI:

```bash
export PULSELOG_DASHBOARD=false
export PULSELOG_AUTO_OPEN=false
```

---

## ⚡ Performance

PulseLog is designed with performance-sensitive workloads in mind. Logging is non-blocking by design — the `log.info()` call returns immediately while a background worker handles delivery.

### At a Glance

| Metric | Value | Notes |
|--------|-------|-------|
| Hot-path latency | ~2 µs | Non-blocking enqueue |
| Single-thread throughput | ~450k/sec | Messages queued |
| 64-thread throughput | ~28M/sec | Near-linear scaling |
| Memory per record | ~80 bytes | Bounded queue |
| Memory leaks | None | Verified 30s sustained load |

### Hot-Path Latency

The key performance property: **logging does not block your application**.

```python
log.info("message")  # Returns in ~2-3 µs (message queued, not delivered)
```

End-to-end latency (queue → handler) depends on `worker_interval`:

| worker_interval | P50 latency | P99 latency |
|-----------------|-------------|-------------|
| 0.001s (1ms)    | ~0.6ms      | ~1ms        |
| 0.01s (10ms)    | ~6ms        | ~7ms        |
| 0.1s (100ms)    | ~60ms       | ~70ms       |

### Throughput

Producer throughput (messages queued per second):

```
Single thread:    ~450,000 ops/sec
8 threads:        ~1,800,000 ops/sec
64 threads:       ~28,000,000 ops/sec
512 threads:      ~220,000,000 ops/sec
```

These numbers reflect the enqueue rate. Actual handler processing depends on handler implementation.

### Concurrency

PulseLog scales well with concurrent producers:

| Threads | Throughput | Degradation vs 1T |
|---------|------------|-------------------|
| 1       | 450k/sec   | baseline          |
| 8       | 1.8M/sec   | none (superlinear)|
| 64      | 28M/sec    | none (superlinear)|
| 128     | 61M/sec    | none (superlinear)|
| 256     | 116M/sec   | none (superlinear)|
| 512     | 220M/sec   | <10% degradation  |

### Memory

PulseLog uses bounded memory by design:

- Queue size is configurable (default: 10,000)
- No unbounded growth
- Zero memory leaks detected in sustained load testing (30+ seconds, 10M+ messages)
- LogRecord overhead: ~80 bytes per message

### Backpressure Behavior

When the queue is full, new messages are dropped and counted:

```
Queue capacity: 1,000
Messages sent:  100,000
Result:         ~38,000 processed, ~62,000 dropped (counted, not lost silently)
```

This is by design — it prevents logging from causing out-of-memory errors in your application. Monitor `stats()["records_dropped"]` if you need to track drops.

To reduce drops under high load:

- Increase `queue_size` (trades memory for drop tolerance)
- Decrease `worker_interval` (trades CPU for faster drain)
- Reduce message volume or batch logs

### Exception Logging

Exception logging (with traceback) is slower due to Python's `traceback.format_exc()`:

```
Plain log.info():           ~450,000 ops/sec
log.exception():           ~75,000 ops/sec  (6x slower)
```

This is expected and acceptable since exception logging should be rare in production.

### What Affects Performance

| Factor | Impact |
|--------|--------|
| `worker_interval` | Lower = lower latency, slightly higher CPU |
| `queue_size` | Larger = more buffering, more memory |
| Message size | Minimal impact (messages are references, not copied) |
| Structured fields | Minimal impact |
| Number of handlers | Linear impact per handler |
| CPU-bound competitors | Significant (GIL contention) |

### Running Benchmarks

```bash
# Quick benchmark
python -m pulselog.benchmark

# Stress test (longer, adversarial scenarios)
python -m pulselog.benchmark --stress
```

Benchmark results depend on Python version, OS, CPU, and configuration. Always run benchmarks in your own environment for accurate numbers.

---

## 🔒 Reliability

PulseLog is designed around:

- bounded buffering
- asynchronous processing
- graceful shutdown
- structured records
- checkpoint persistence
- concurrent producer support
- configurable runtime behaviour

Applications should still treat logging as an auxiliary system and avoid placing critical business state exclusively in logs.

---

## 🧩 Design Goals

| Goal | Description |
|---|---|
| Low application overhead | Keep logging work away from the main application path |
| Non-blocking operation | Avoid waiting for dashboard consumers |
| Bounded memory | Prevent an unlimited logging backlog |
| Structured data | Preserve useful metadata |
| Concurrency | Support multiple producer threads |
| Batch processing | Process pending records efficiently |
| Real-time visibility | Make application behaviour visible in a browser |
| Resumable workflows | Provide checkpoint support |
| Python-first API | Keep the public API simple |

---

## 📦 Requirements

- Python 3.8+
- `websockets >= 11.0`

For Python versions below 3.11, PulseLog uses `tomli` for TOML configuration support.

---

## 🧪 Development

```bash
git clone <repository-url>
cd pulselog

pip install -e ".[dev]"
```

Run tests:

```bash
pytest
```

---

## 📊 Benchmarking

The project includes a dedicated benchmark suite covering:

```text
Latency
Throughput
Concurrency
CPU
Memory
Queue pressure
Drops
Shutdown
Sustained load
```

Recommended concurrency levels:

```text
1
2
4
8
16
32
64
```

Recommended message sizes:

```text
32 B
128 B
512 B
1 KB
4 KB
16 KB
64 KB
```

Recommended structured-field counts:

```text
0
1
5
10
25
50
100
```

Benchmark results should always include the machine and Python environment used for the measurement.

---

## 🧭 Roadmap

PulseLog is evolving toward a production-grade logging and observability tool for Python workloads.

Potential areas of development include:

- further performance improvements
- improved concurrent logging
- richer structured logging
- stronger backpressure controls
- improved runtime metrics
- dashboard improvements
- persistent log storage
- OpenTelemetry integration
- distributed logging support
- additional integrations
- expanded platform support
- long-duration stability testing

---

## 📌 Project Status

Current version:

```text
2.0.0
```

Development status:

```text
Alpha
```

PulseLog is actively evolving and the public API may change during early releases.

For reproducible deployments:

```bash
pip install "pulselog==2.0.0"
```

---

## 📄 License

MIT License.

---

<div align="center">

**⚡ PulseLog**

*Keep your application moving.*

</div>


