Metadata-Version: 2.4
Name: ollambda
Version: 0.1.4
Summary: Add your description here
Requires-Python: >=3.13
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: configlite>=0.2.5
Requires-Dist: flask>=3.1.3
Requires-Dist: gunicorn>=26.0.0
Provides-Extra: test
Requires-Dist: pytest>=9.1.1; extra == "test"
Requires-Dist: pytest-cov>=7.1.0; extra == "test"
Requires-Dist: pre-commit>=4.6.0; extra == "test"
Dynamic: license-file

# ollambda

Simple Flask based API server to proxy ollama calls to a more performant backend (such as llama.cpp)

## Why should I use this package?

Honestly, you probably shouldn't. Not only is it currently nowhere near completion, there are plenty of alternatives.

Ollambda mostly exists as a personal learning tool.

But if you still want to, then continue on to see usage instructions and more info.

If you find a bug, feel free to open an issue, or make a pull request. Contributions are welcome!

## How to Use

### Installation

To install ollambda, either install using `pip install ollambda` or clone this repository and use `pip install .`

#### venv

Venv usage is encouraged, the simplest way to manage this is to use [uv](https://docs.astral.sh/uv/):

```
cd ollambda
uv venv .venv
source .venv/bin/activate
uv pip install .
```

### Configuration

Currently ollambda proxies requests to an upstream llama.cpp server. The address is configured via environment variables or a YAML config file at `~/.ollambda/config.yaml`:

| Setting    | Env var               | Config key   | Default            |
| ---------- | --------------------- | ------------ | ------------------ |
| Proxy IP   | `OLLAMBDA_PROXY_IP`   | `proxy_ip`   | `127.0.0.1`        |
| Proxy port | `OLLAMBDA_PROXY_PORT` | `proxy_port` | `8080`             |
| Log Dir    | `OLLAMBDA_LOG_DIR`    | `log_dir`    | `~/.ollambda/logs` |

Ollambda uses [ConfigLite](https://github.com/ljbeal/ConfigLite) for its config backend, which I am also the developer of.

### Running the server

Start with the bundled command (uses Gunicorn internally):

```bash
ollambda
```

This spawns one worker per CPU core and binds to `0.0.0.0:5000` by default.

For manual control, you can also invoke Gunicorn directly:

```bash
gunicorn ollambda:app -w 4 -b 0.0.0.0:5000
```

## ollama API coverage

| Method | Endpoint             | Description                                                 | Implemented? |
| ------ | -------------------- | ----------------------------------------------------------- | ------------ |
| POST   | `/api/generate`      | Generate a completion (streaming or non-streaming)          | N            |
| POST   | `/api/chat`          | Generate a chat completion (with tool calling support)      | Y            |
| POST   | `/api/create`        | Create a model from another model, GGUF, or safetensors     | N (no plans) |
| GET    | `/api/tags`          | List locally available models                               | Y            |
| POST   | `/api/show`          | Show model information (details, modelfile, template, etc.) | Y            |
| POST   | `/api/copy`          | Copy a model to a new name                                  | N            |
| DELETE | `/api/delete`        | Delete a model                                              | Y            |
| POST   | `/api/pull`          | Download a model from the Ollama library                    | N            |
| POST   | `/api/push`          | Upload a model to a model library                           | N (no plans) |
| POST   | `/api/embed`         | Generate embeddings from a model                            | Y            |
| GET    | `/api/ps`            | List models currently loaded in memory                      | Y            |
| GET    | `/api/version`       | Retrieve the Ollama server version                          | Y            |
| HEAD   | `/api/blobs/:digest` | Check if a file blob exists on the server                   | N (no plans) |
| POST   | `/api/blobs/:digest` | Push a file blob to the server                              | N (no plans) |

> [!IMPORTANT]
> Current "implementation" may not be fully complete and API aligned.

> [!NOTE]
> Source: https://github.com/ollama/ollama/blob/main/docs/api.md
