Metadata-Version: 2.4
Name: uring-api
Version: 0.1.0rc5
Summary: Small Python wrapper around Linux io_uring
Author-email: Kristjan Valur Jonsson <sweskman@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/kristjanvalur/pytealet/tree/main/packages/uring_api
Project-URL: Repository, https://github.com/kristjanvalur/pytealet
Project-URL: Source, https://github.com/kristjanvalur/pytealet/tree/main/packages/uring_api
Project-URL: Changelog, https://github.com/kristjanvalur/pytealet/blob/main/packages/uring_api/CHANGELOG.md
Project-URL: Issues, https://github.com/kristjanvalur/pytealet/issues
Keywords: io_uring,linux,async,proactor
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: C
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: 3.15
Requires-Python: >=3.10
Description-Content-Type: text/markdown

# uring-api

`uring-api` is a small Python wrapper around Linux `io_uring`.

The goal is deliberately modest: expose enough of the native ring lifecycle,
socket send/recv submission, completion waiting, and callback delivery to build
higher-level completion abstractions in Python. It does not implement an event
loop, scheduler, or asyncio compatibility layer.

Future work is tracked in [ROADMAP.md](ROADMAP.md), including queue resizing
and specialised kernel tuning. Caller-owned provided-buffer receive with leased
`BufView` delivery is already part of the Python surface.

## Quick Check

```python
import uring_api

print(uring_api.probe())

with uring_api.Ring() as ring:
    print(ring.fd)
```

## Socket I/O

Need to drive socket work through a ring without building a full event loop?
`Ring` exposes direct prepare wrappers for the common Python-oriented cases:

- stream I/O: `prepare_recv()`, and provided-buffer `prepare_recv_buf()` /
  `prepare_recv_multishot()` via `create_buf_group()`;
- message I/O: `prepare_recvmsg()`, `prepare_sendto()`, `prepare_sendmsg()`, and
  zero-copy `prepare_sendmsg_zc()`;
- listeners and setup: `prepare_accept()`, `prepare_accept_multishot()`,
  `prepare_connect()`, and `prepare_socket()`;
- lifecycle: `prepare_shutdown()`, `prepare_close()`, nowait
  `prepare_*_nowait()` for close/shutdown/cancel/poll_remove, send helpers
  `prepare_send()` / `prepare_send_zc()`, `wait()` for completion reaping, and
  `poll()` to wait until the CQ is non-empty without harvesting.

Each submitted operation carries a Python `user_data` object which comes back
with its completion. `completion.take_user_data()` returns that payload and
clears the slot — the usual completion-callback pattern when a waitable
reverse-links the `Completion`. `completion.user_data` is still **settable**
(assign `None` or `del` to clear without taking). Kernel SQE identity remains
the `Completion` pointer, not `user_data`. Inspect the semantic operation with
`completion.kind` (`CompletionKind` enum, or the matching `COMPLETION_KIND_*`
constants) when callbacks need to branch on completion type rather than
inferring from `result` alone.

### Multishot delivery contract

Multishot ops (`prepare_*_multishot`) keep one **armed** `Completion` handle —
the object returned from prepare, also the cancel / poll_remove target. Kernel
SQE `user_data` always points at that handle.

- **Intermediate legs** (`IORING_CQE_F_MORE`): delivery is a fresh **shell**
  `Completion` that copies `user_data` (and leg `sequence`) from the armed
  handle. `take_user_data()` on the armed handle defers while more CQEs are
  staged, so a concurrent `!MORE` delivery cannot clear the slot first. The
  armed object is left untouched; shells do not re-arm reverse links.
- **Terminal leg** (`!MORE`, including cancel / poll_remove / `-ENOBUFS` /
  stream end): delivery **is** the armed handle itself. Call
  `completion.take_user_data()` (or assign `None`) to drop the cycle with
  any waitable that stored the reverse link.

Clients that reverse-link waitable → `Completion` should take `user_data` on
every delivered object (shell or terminal); only the terminal call hits the
armed handle, and that clear waits until every staged leg has been packaged.
Do not assume every multishot CQE is a distinct object — only MORE legs are.

**In-flight handle.** Prepare holds an extra reference on the armed
`Completion` because the kernel SQE stores only a pointer. Oneshot drops that
ref when its CQE is packaged. Multishot and zero-copy send share the pointer
across several CQEs, so the extra ref stays until every **already staged** CQE
for that handle has been turned into a Python object (MORE shells first, then
the parent on `!MORE`; for `send_zc`, until the internal `NOTIF` CQE). A
per-handle counter (`aux_refcount`) plus a sticky “terminal was staged” flag
(`AUX_DECREF`) make it safe to package `!MORE` before an earlier MORE row in
the same batch — the parent is not `DECREF`’d while a shell still needs it.

**Safe clear.** `take_user_data()` on a shell or idle handle steals the
payload and writes `None` immediately. On an armed handle that still has
staged CQEs it returns a new reference, sets `USER_DATA_CLEAR`, and leaves
the live slot so a not-yet-built MORE shell can copy the waitable. The real
clear runs after the last staged leg is packaged (same window as the
in-flight ref). The flag and the pointer are updated together under the
ring’s refcount mutex; the old waitable is released after the lock is
dropped. Assigning `None` or `del` is the same deferred-clear window without
returning the payload. Assigning any other object replaces the slot
immediately, including while a MORE shell is being built. The shell copies
the pointer under that same mutex and keeps its own reference, so it sees
either the previous object or the new one. A shell that already exists keeps
the object it copied. The terminal leg is the armed handle: its callback
reads whatever is in the slot when it loads `user_data`.

Calling `take_user_data()` or assigning `user_data` after the `Ring` object
has been deallocated is **undefined**. `ring.close()` is fine (the mutex
still belongs to the live `Ring`); dropping the last reference so the ring
is collected while a handle remains is not a supported use.

**Send-all:** `prepare_send_all(fd, data)` (or `construct_send_all` then
`prepare`) is a synthetic drain: the kernel still sees ordinary send SQEs, but
Python gets one `Completion` when the buffer is exhausted. Partial CQEs re-arm
the remainder internally (`POLL_FIRST` on later legs when probed).
`Completion.timeout` set before `prepare` is per leg, not a deadline for the
drain. Each submitted send is linked to a new relative timer of that full
duration. Time spent parked, or between a partial completion and the next
submit, does not count, and a peer that keeps accepting data can outlast
`timeout`. A leg that stalls finishes the drain with `-ECANCELED` or
`-EINTR`. If a later leg cannot allocate its timer, the bytes already
accepted stay parked and `wait()` does not fail. A later drain that still
cannot allocate leaves that leg queued, and neither `wait()` nor `submit()`
fails. The drain fails only when that park cannot be queued. Success `res`
is the total byte count, clamped to `INT_MAX`;
`result` is the full unsigned count. Zero-byte send on a non-empty remainder fails
with `-EAGAIN`. `skip_success` keeps successful
drains off `wait()` / `callback` and delivers the handle on failure.
`skip_all` skips user delivery entirely (errors use
`nowait_error_handler`). Unlike tagged nowait helpers, `send_all` still
holds the prepare in-flight ref and is included in `pending_count()` until the
drain terminals. `prepare_cancel` of the handle abandons further legs: a parked
continuation completes `-ECANCELED` instead of flushing another send.
`Completion.no_deliver_multi` only affects `recv_multishot`. Set it once the
caller wants nothing more from the connection, so later CQEs of that receive
are not delivered through `wait()` or a callback: MORE legs, EOF (`res == 0`),
other errors, and the terminal `-ECANCELED`. That avoids an extra completion
for each leftover chunk. Accept multishot, poll multishot, and oneshot
completions are still delivered, flag or not. The flag may be set on those
kinds; it is ignored. Buffers and the in-flight ref are released first. An
omitted MORE leg does not allocate a shell `Completion`. An omitted terminal
leg keeps `res` and `flags` on the armed handle and does not allocate a
`BufView`, including the empty EOF view. The flag is not copied onto MORE
shells (the check reads the armed handle), and unlike `skip_success` it can
be set after `prepare`.
`prepare_cancel(..., no_deliver_multi=True)` and the matching
`construct_cancel` / `*_nowait` helpers set that flag on the **target**
before the cancel SQE is submitted. That is a convenience, not a request to
hide the cancel completion: the waitable cancel is still delivered. Later
`recv_multishot` CQEs are dropped even if that cancel never enters the kernel. If
that call fails before a cancel SQE exists, a bit it just set is cleared, so
a later CQE is not swallowed. `no_deliver_multi=False` does not clear a flag
you set yourself, and a failed call does not clear that either.
A packer that fills a next-leg `io_uring_submit`s it when this thread may
enter and a unique waiter is already held (it may be blocked in
`wait_cqe`). Otherwise the SQE stays
prepared until the next harvest flush or host `submit()` / `wait()`, or it
parks on fill-wait when there is no slot. With `auto_submit` off, `wait()` does not
publish them — call `submit()` as with any other prepared SQE (including after
an empty wait batch while `pending_count()` is still non-zero). The next user
`prepare` fills parked next-legs first (they take the SQ slot ahead of the new
op; `auto_submit` makes room if the SQ is full). `submit()` never raises
`SubmissionQueueFull`: a full SQ is submitted first, then parked legs are
filled. The returned count is every SQE that enter submitted, including those
flushed to make room.
While a send-all is busy on an fd,
`prepare` of send/close/shutdown/another send-all on that fd parks on a
per-fd conflict FIFO (`prepared` stays false until drain copies it into the
SQ). Recv is full-duplex and still fills an SQE. `sendto` is datagram and
does not park (it is not mixed with stream send-all); `sendmsg` on a stream
still conflicts. `prepare([send_all, close])`
therefore serialises in one batch. `prepare()` returns the number **accepted**
(SQE fills, conflict-FIFO parks, and fill-wait parks); `Completion.prepared`
is true only after an SQE is filled, so a parked close stays `prepared is
False`. A conflicting `prepare` parks on the FIFO before leftover drain, so a
full SQ does not raise for send/close on a busy fd. Recv and other fds still
drain leftovers (fill-wait, then FIFO) first so parked next-legs take the next
SQ slot. CQE drain fills SQEs while a slot exists; `io_uring_enter` only when
`auto_submit` is on and this thread may submit. Cancel of the active drain still fills an SQE; cancel of a
**queued** op stays behind it; cancel of an already-prepared (SQ / in-kernel)
send on that fd fills an SQE now. Issuer `auto_submit=False` still raises
`SubmissionQueueFull` from `prepare` — it does not spill onto the conflict
FIFO. A non-issuer that would have to enter parks on the fill-wait list.
Once an fd has used send-all, later send/shutdown/close on it should go
through the ring until that fd is idle (libc `close()` while a drain is live
stales the table).

**Where a Completion sits after `prepare`:**

| Place | Role | `prepared` |
| --- | --- | --- |
| Kernel SQ | Lazy batch. Next `submit()` / wait flush publishes it. | true |
| Conflict FIFO | Per-fd busy serialisation behind send-all. | false |
| Fill-wait | This thread cannot `io_uring_enter` (SQ full, non-issuer). | false |

Drain order is **fill-wait, then that fd's conflict FIFO, then the new op**.
`prepare` is accepted in all three seats; `Completion.prepared` is true only
after an SQE fill. Issuer `auto_submit=False` still raises `SubmissionQueueFull`
instead of parking on fill-wait.

**Lazy submit:** `prepare_*` / nowait helpers (including cancel and poll_remove)
only fill SQEs. Work becomes kernel-visible when you call `ring.submit()`,
when **`auto_submit` is on (the default) and `wait()` flushes a submission queue
with `sq_waitable` set** (if this thread may submit), when prepare hits a full SQ, or after an
inline `wait()` delivery batch. A cancel sets `sq_waitable`, so the next
`wait()` submits it, unless that cancel is nowait **and** its target is a
`recv_multishot` with `no_deliver_multi` set: neither the cancel ack nor the
target CQE is delivered, so there is nothing to wake for. A direct close,
shutdown, or poll_remove does not set the bit. The same op copied out of a
park (behind `send_all`, or fill-wait) does: the wait that releases that tail
submits it. A queue of only direct non-waitable SQEs is not flushed
by `wait()` or `serve_completions()`. Call `submit()`, or prepare a waitable
SQE and let the next flush take both: the queue is ordered, so a non-waitable
SQE already ahead of that waitable one rides the same enter. The unique CQ waiter in `serve_completions()` uses the same rule before harvest
when this thread may enter. TAKE workers never `io_uring_submit` — submitting
after every CQE unbatches the SQ against a driving thread. Set
`Ring(..., auto_submit=False)` or `ring.auto_submit = False` so the **issuer**
raises `SubmissionQueueFull` instead of flushing from prepare, and so `wait()`
leaves prepared SQEs unsubmitted until you call `submit()`. A non-issuer `prepare` that
would have to enter parks on the fill-wait list. `Ring.prepare(...)` returns
the number of entries successfully prepared (SQE fills and parks). With
`auto_submit` on, do not call `submit()` before every `wait()` — wait does
that. With completion workers parked only on `wait_idle`, the issuer still
flushes before that park (workers never call `wait()`).

A blocking or timed `wait()` that reaps only silent CQEs (nothing delivered,
and not a `break_wait`) submits any waitable SQEs prepared while handling that burst
and parks again, when this thread may submit. A timed wait keeps the original
deadline and parks with the time still left. `wait(0)` returns after one
harvest. `break_wait` still returns: a wake NOP, or a sticky latch taken
before the reaper entered the kernel, is not retried.

**Pending count:** `ring.pending_count()` is waitable `Completion`s that still
hold the prepare in-flight ref, plus one for each link-timeout submission
whose timer completion has not been consumed. The completion count goes up at
successful waitable `prepare` (SQE fill, conflict-FIFO enqueue, or fill-wait
enqueue), and down when that ref is dropped (oneshot CQE packaged, or
multishot / `send_zc` / `send_all` after the terminal CQE). Each filled
link-timeout SQE adds one until that CQE is consumed and its timespec is
freed. The timer is not delivered. It is often in the same harvest as the
operation; when it is not, the count stays non-zero so a drain keeps waiting.
Closing while the count is still non-zero abandons the copy, the same as any
other undrained submission. Construct without prepare, ordinary nowait
helpers, and MORE shells do not change the count, unless the nowait op has a
link timeout. Nowait `send_all` keeps the in-flight ref until the drain
terminals.

**Runtime counters:** `ring.stats()` is how full the queues get and who
flushes them. The dict is monotonic — subtract two calls; there is no reset.

| Key | Counts |
| --- | --- |
| `sqe` | SQEs obtained, including the occasional wake NOP |
| `cqe` | CQEs consumed, including NOPs, nowait, link-timeout timers, multishot legs, and zero-copy notifications |
| `sq_full` | A fill attempt's first peek found no free slot |
| `next_leg` | Send-all continuation sends filled (not an abandon NOP) |
| `next_leg_park` | Continuations that could not take a slot and parked on fill-wait |
| `submit_main_events` / `submit_main_sqes` | Owning thread: `submit()`, the `wait()` flush, and the `submit()` before `wait_idle` |
| `submit_worker_events` / `submit_worker_sqes` | `serve_completions()` harvest flush, and a non-owner `break_wait` NOP |
| `submit_next_events` / `submit_next_sqes` | A deliberate send-all continuation enter |
| `submit_sq_full_events` / `submit_sq_full_sqes` | A make-room flush because `get_sqe` found no free slot |
| `wait_calls` | Every `Ring.wait()` that reached the reap, including an empty return |
| `wait_front_events` / `wait_front_cqes` | `Ring.wait()` harvests that got a completion |
| `wait_back_events` / `wait_back_cqes` | `serve_completions()` reaper harvests |
| `cq_overflow` | Kernel overflow count at this call; `0` after `close()` |

`submit_*_sqes / submit_*_events` is the average batch published per enter.
`sqe` is not that sum: the difference is SQEs still sitting in the submission
queue. A wait event is one reap that returned a completion; `*_cqes` is how
many that drain took, so `cqes / events` is how many were gathered at a time.
An empty `wait()` is not an event, so `wait_front_events / wait_calls` is
how often the loop woke with something. `poll()` and `serve_completions()`
are not `wait_calls`. `cqe` is
`wait_front_cqes + wait_back_cqes` (one completion of sampling skew aside).
`cqe` can still run ahead of the submission-side fields.

**Construct then prepare:** every waitable op has `construct_*` (bind cargo,
no SQE) and `prepare_*` (construct + prepare of one handle). Cargo lives on
the matching sidecar; `completion.prepared` is false until an SQE is filled
(a conflict-FIFO or fill-wait park is accepted by `prepare()` but stays
`prepared is False`). Arm a reverse link on the constructed object, then
`ring.prepare(completion)` or `ring.prepare([c1, c2, ...])`. `prepare` returns
the number accepted (SQE fills and parks) and does not submit;
`wait()` / `submit()` flush as usual (or a full SQ if `auto_submit` is on).
On prepare error, earlier entries in the list may already be accepted.

```python
pending = []
for chunk in outgoing:
    completion = ring.construct_send(fd, chunk, 0, token)
    # arm reverse here — nothing can complete yet
    pending.append(completion)
ring.prepare(pending)
batch = ring.wait(1.0)
```

Internal `break_wait` NOPs use tagged `user_data` (wake `…01`). Nowait prepare
stamps a tagged nowait token (`…11` with kind/fd payload), not a `Completion*`;
a constructed nowait handle is only a temporary hold for `prepare` and is never
delivered to the client.

```python
import socket
import uring_api

reader, writer = socket.socketpair()
try:
    reader.setblocking(False)
    writer.setblocking(False)

    with uring_api.Ring() as ring:
        token = {"operation": "greeting"}
        buf = bytearray(5)
        ring.prepare_recv(reader.fileno(), buf, 0, token)
        # optional explicit flush; wait() also flushes first
        ring.submit()
        writer.send(b"hello")

        assert ring.poll(1.0) is True
        batch = ring.wait(0)

    assert len(batch) == 1
    completion = batch[0]
    assert completion.user_data is token
    assert bytes(buf) == b"hello"
    print(completion.res, completion.result)
finally:
    reader.close()
    writer.close()
```

For sends, `uring-api` keeps the exported buffer alive until the kernel reports
the completion. That avoids copying the outgoing payload into an internal bytes
object just to keep memory valid. `prepare_send_zc()` uses
`IORING_OP_SEND_ZC`, while `prepare_sendmsg_zc()` uses `IORING_OP_SENDMSG_ZC` for
the `sendmsg` shape. Their ordinary operation CQE is delivered as the submitted
`Completion`; the later `IORING_CQE_F_NOTIF` buffer-lifetime CQE is consumed
internally and releases the retained buffer.

`prepare_shutdown()` is a socket operation and mirrors `shutdown(fd, how)`.
`prepare_*` / `construct_*` take SQE cargo first and `user_data` last.
METH_FASTCALL helpers (send, accept, poll, close, …) are positional-only: a
three-arg `prepare_send(fd, data, x)` is flags, not a token. `openat` is
`prepare_openat(dfd, path, flags, mode=0, user_data=None)`.
`prepare_recv` / `prepare_recvmsg` take `flags` like send. Include
`IORING_RECVSEND_POLL_FIRST` to poll before the first recv/send; the
extension puts that bit in SQE `ioprio` and leaves other `MSG_*` in
`msg_flags`. **`POLL_FIRST` is `1 << 0`, the same bit as `MSG_OOB`:**
that value is treated as poll-first, not out-of-band. Do not pass
`socket.MSG_OOB` in this word. A separate ioprio argument can wait
until we need more `IORING_RECVSEND_*` bits.
`probe()["IORING_RECVSEND_POLL_FIRST"]` is the 5.19 floor. Do **not**
combine `POLL_FIRST` with `prepare_recv_multishot` — liburing does not
test that pairing, and the kernel can leave a `MORE` handle with no
terminal CQE. The bit is ignored on multishot prepare.
`IORING_CQE_F_SOCK_NONEMPTY` may appear on `completion.flags` after a recv
when more data is already queued.
`prepare_accept()` and `prepare_accept_multishot()` accept optional accept flags;
pass `socket.SOCK_NONBLOCK | socket.SOCK_CLOEXEC` when accepted sockets should
be ready for proactor ownership without a follow-up `fcntl()` call.
Multishot first-leg numbering is `completion.sequence`. Python
`construct_*_multishot` / `prepare_*_multishot` take an optional last
positional `base_sequence=0` after `user_data` so the index is set before
the SQE is filled (needed when auto_submit or SQPOLL can complete before
the caller runs another statement). `prepare_recv_multishot` is
`fd, buf_group, flags=0, user_data=None, base_sequence=0`. Setting
`completion.sequence` after construct still works; do not set it after
`prepare_*` returns. C construct stays cargo-only (`completion_set_sequence`
after construct).
`prepare_accept()` and `prepare_accept_multishot()` deliver the accepted fd in
`completion.res` and `completion.result`. Call `getpeername()` on the fd when
you need the peer address.
`prepare_close()` is lower-level: pass only a raw fd whose ownership has already
been transferred away from Python objects such as `socket.socket`, for example
with `detach()`. Otherwise, Python and the kernel may both believe they own the
same descriptor. When you do not need a result or waitable handle, nowait
helpers construct a temporary `Completion`, prepare a tagged nowait SQE, and
drop the handle: `prepare_close_nowait(fd)`,
`prepare_shutdown_nowait(fd, how)`, `prepare_cancel_nowait(completion)`, and
`prepare_poll_remove_nowait(completion)`. They return `None`, and never deliver via `wait()` or callbacks.
`prepare_cancel` and `prepare_cancel_nowait` take keyword-only
`no_deliver_multi=False`. When set, the target's `no_deliver_multi` flag is
stored before the cancel SQE is posted. On a `recv_multishot` target, further
CQEs (MORE legs, EOF, errors, and the terminal `-ECANCELED`) are consumed in C
and do not show up in `wait()`. Accept, poll, and oneshot targets are not
affected. The cancel does not have to enter the kernel for that to take
effect. The cancel request itself, when you used the waitable helper, is
still delivered.
To batch with waitable ops, use `construct_close_nowait(fd)` (or set
`completion.skip_all = True` on a constructed close/shutdown/cancel/poll_remove)
and pass it to `prepare`. On kernels with
`IORING_FEAT_CQE_SKIP`, successful nowait ops post no CQE
(`IOSQE_CQE_SKIP_SUCCESS`). Failed nowait CQEs (`res < 0`) invoke optional
`Ring.nowait_error_handler` (successful CQEs, when posted without
`CQE_SKIP_SUCCESS`, are dropped silently) with a context dict (`message`,
`ring`, `res`, `flags`, `kind` as `COMPLETION_KIND_*`, and advisory `fd` or
`None`). `Completion.skip_all` is that fire-and-forget flag (implies
`skip_success`). `Completion.skip_success` alone keeps the handle, skips
successful delivery, and completes **only on error** (no kernel
`CQE_SKIP_SUCCESS`; counted in `pending_count()` until the CQE). `send_all`
always keeps the handle. `user_data` is only a token.
Nowait cancel acks of `-ENOENT` / `-EALREADY` (target already gone
or already completing) are dropped silently; waitable cancel still reports
those as `res < 0`. If that hook raises, `exception_handler` is used; the
CQ drain always continues.

## File metadata and positioned I/O

`Ring` also exposes positioned file helpers for caller-owned fds:

- `prepare_openat(dfd, path, flags, mode=0, user_data=None)` opens a path and
  returns the new fd in the completion result (`dfd=AT_FDCWD` for a cwd-relative
  path);
- `prepare_read(fd, buf, offset)` and `prepare_write(fd, data, offset)` perform
  explicit-offset I/O into caller buffers;
- `prepare_statx_fdsize(fd)` is the common fast path for open-file metadata: it
  runs fd-only statx internally and puts the byte length in `completion.result`
  on success (`completion.kind == CompletionKind.STATX_FDSIZE`);
- `prepare_statx(dfd, path, flags, mask, buf)` fills a caller-provided 256-byte
  statx buffer asynchronously when you need a custom mask or path lookup.

The usual positioned-file case (append EOF, `SEEK_END`, sendfile bounds) is an
open fd whose size you already own:

```python
handle = ring.prepare_statx_fdsize(fd)
[completion] = ring.wait()
if completion.res == 0:
    size = completion.result
```

No caller buffer is required for `prepare_statx_fdsize()`. Use `prepare_statx()`
when you need path-based metadata or fields beyond `stx_size`.

Successful `prepare_statx()` completions always leave `completion.result` as
`None`; read fields from the caller-owned submit buffer (for example via
`statx_st_size(buf)` when you requested `STATX_SIZE`). Only
`prepare_statx_fdsize()` puts the byte length in `completion.result`. If the
internal buffer lacks size fields, `completion.result` is `None` and the
completion is still delivered.

**Behaviour change (since PR #34):** successful `prepare_statx()` no longer sets
`completion.result` to `0`; it stays `None`.

Provided-buffer receive uses a caller-owned ring created with
`create_buf_group()`. Submit one-shot receives with `prepare_recv_buf()` or
stream receives with `prepare_recv_multishot(fd, buf_group, ...)`. Both paths
return read-only `BufView` objects rather than copying into `bytes`. Export the
payload with `memoryview(view)` and drop the export (or call
`memoryview.release()`) before the kernel buffer is recycled:

```python
buf_group = ring.create_buf_group(buffer_size=16384, buffer_count=256)
pending = ring.prepare_recv_buf(reader.fileno(), buf_group, 0, token)
[completion] = ring.wait(1.0)

view = memoryview(completion.result)
try:
    process(view)
finally:
    del view
```

Multishot receive reuses the same `BufGroup` contract. Each CQE delivers a
leased `BufView`, sets `completion.sequence` for out-of-order callback
reconstruction, and uses `IORING_CQE_F_MORE` until EOF, cancellation, or
`-ENOBUFS` when the buffer ring is empty. MORE legs are shell Completions;
the terminal `!MORE` is the armed submit handle (see multishot delivery
contract above). After `-ENOBUFS`, return buffers to the ring and submit a
fresh `prepare_recv_multishot()`; stream consumers should continue ordinal
indexing from the terminal completion's `sequence`.

```python
handle = ring.prepare_recv_multishot(reader.fileno(), buf_group, 0, token)
[completion] = ring.wait(1.0)
view = memoryview(completion.result)
try:
    process(view)
finally:
    del view
```

`BufView` tracks active exported memoryviews and recycles the selected buffer
back to the ring when the last export is released. Provided-buffer completions
always return `BufView`, including EOF (`completion.res == 0`), where the view
has `length == 0` and is falsy. Detect stream end from `completion.res`, not
from the result type. `BufGroup` and `BufView` cannot be constructed directly;
use `Ring.create_buf_group()` and let receive completions create the views.

`BufGroup` supports an optional **owner hook** for reusing pools without
wrappers:

- Set `buf_group.release_callback = callable` (or `None`).
- `buf_group.inflight_count` is how many `recv_buf` / `recv_multishot`
  SQEs are still armed on this group. It goes up when the SQE is filled
  (not at construct, and not while the op is only parked on a full queue)
  and down once on the terminal `!MORE` CQE, including when that CQE is
  not delivered. It is not `leased_count`: an armed recv that has not
  selected a buffer holds no view. `buf_group.in_use()` is true when
  either count is non-zero.
- With `release_callback` set, `close()` does **not** unregister. The first
  call sets `release_invoked` and calls the hook on the calling thread,
  whether or not a receive is armed. A later `close()` does not call the
  hook again until the owner sets `release_invoked` false (a cache does
  that when it hands the group out). If the hook raises, the flag stays
  set; clear it before retrying. The hook is not deferred until the
  group goes idle: the terminal completion may be reaped on another thread,
  and uring-api does not marshal it back. The owner keeps the group (for
  example a size-keyed cache parking it for the next checkout). Clear the
  hook before a real dispose. `close()` is not thread-safe, and this path
  does not take a lock. `close_mu` only pairs a hard release with the
  terminal completion's inflight drop.
- With no hook, `close()` is a hard release. Idle groups unregister
  immediately. An armed group stays registered until the last terminal
  CQE: unregistering that bgid and handing it to a new group lets the old
  request write into the new storage. A second hard `close()` before that
  CQE does nothing.
- Do not arm a new receive on a group you have hard-closed. `close()` does
  not interlock with `prepare`.
- Finalization still frees the group if nothing called `close()`; dealloc does
  **not** call `release_callback` (abandoned groups are not returned to a
  cache). Clear the hook before a real dispose so `close()` destroys the
  group rather than re-entering the owner.

```python
free: list[uring_api.BufGroup] = []


def return_to_cache(group: uring_api.BufGroup) -> None:
    free.append(group)


group = ring.create_buf_group(16384, 4)
group.release_callback = return_to_cache
group.close()  # idle: returns to free list; ring buffers stay registered
assert free == [group]

group.release_callback = None
group.close()  # frees the provided-buffer ring
```

`completion.kind` uses `RECV_MULTISHOT` (13) for multishot provided-buffer
receive and `RECV_BUF` (16) for one-shot `prepare_recv_buf()`.

`CompletionKind` values are stable across releases and mirror the C API
constants in `uring_api_completion_kinds.h`. Prefer the enum in Python code;
native clients include the same header and call `completion_kind()` on the
completion object.

The local liburing headers expose more socket-adjacent operations than this
wrapper publishes, but those are intentionally outside the core Python-oriented
surface. Readiness polling is optional for a completion proactor, fixed-buffer
send variants still need a different ownership contract than leased `BufView`
receive, and socket command or NAPI controls are specialised tuning hooks. Those
items are tracked in [ROADMAP.md](ROADMAP.md) rather than implied by `probe()`,
which remains a compact runtime availability check.

When the SQ is full, prepare paths flush pending entries and retry. With
`IORING_SETUP_SQPOLL`, after a second flush that still leaves the queue
completely full, they wait for the kernel poller to free a slot and retry
(not a CQE wait). If some slots are free but fewer than this prepare needs,
that wait would return immediately, so the same path sleeps and retries until
enough slots are free or the deadline passes. Non-SQPOLL rings must free a
slot after one successful flush. If the slots still cannot be obtained,
prepare raises `RuntimeError` — a stuck queue or dead poller, not ordinary
backpressure. A link timeout asks for two free slots through that same path
and does not take either slot until both are free. If `sq_entries` is less
than two, that prepare raises `RuntimeError` immediately: the queue cannot
hold the pair, so this is not a dead poller and nothing is submitted.
`SubmissionQueueFull` is only the case where this call was not allowed to
enter (`auto_submit` off, or not the submit thread) and the ring is large
enough to hold the request.

## Checking Availability

`io_uring` availability depends on more than the Python package importing
successfully. The kernel, container sandbox, seccomp profile, and process limits
can all affect whether a ring can actually be created.

Use `probe()` when you want a compact availability and capability dictionary:

```python
import uring_api

probe = uring_api.probe()

if probe:
    print("io_uring is available")
    print("capabilities:", probe)
else:
    print("io_uring is not available")
```

Use `is_available()` when you only need a boolean:

```python
import uring_api

if not uring_api.is_available():
    raise RuntimeError("io_uring is not available in this environment")
```

`probe()` creates a tiny temporary ring and closes it right away to test ring
creation with the requested `entries` and `flags`. Targeted capability probes run
once per process and are cached in static variables; later `probe()` calls reuse
those results. If ring creation fails,
it returns an empty dictionary. If it succeeds, the dictionary contains
`"available": True` plus named optional capabilities such as
`"IORING_ACCEPT_MULTISHOT"`, `"IORING_POLL_MULTISHOT"`, `"IORING_RECV_MULTISHOT"`, and
`"IORING_OP_SEND_ZC"` and `"IORING_OP_SENDMSG_ZC"` (version-gated at kernel
6.0 per `io_uring_enter(2)`), and `"IORING_OP_STATX"` (version-gated at 5.6).
Production code should
still handle `OSError` when it creates the real ring because limits or sandbox
policy may differ for larger settings.

Pass setup flags to `probe(flags=...)` to check whether this build and kernel
combination accepts a ring mode before using it for the real ring:

```python
import uring_api

flags = uring_api.IORING_SETUP_SINGLE_ISSUER
probe = uring_api.probe(flags=flags)

if probe:
    print("setup flags accepted")
else:
    print("setup flags rejected")
```

Some flags also impose application-level contracts. For example,
`IORING_SETUP_SINGLE_ISSUER` means callers must **submit** (`io_uring_enter`)
from a single owning thread even on kernels that accept the flag. The owner
is the thread that **created** the ring (same as the kernel). Filling an
SQE is not that: any thread may `prepare` if the SQ has a slot. A non-issuer
that would have to enter parks on a fill-wait list until the issuer
`submit()` / `wait()` copies it into the SQ. Send-all next-leg uses the same
list. `submit()` from a non-owner still raises. Construct the ring on the
event-loop thread; do not create it on a factory thread and hand it over.

**Caveat:** `SINGLE_ISSUER` plus `serve_completions` workers. A worker may fill
a send-all next-leg but cannot enter. That SQE stays prepared until the owning
thread `submit()`s (or `wait()` flushes). If the driver parks forever in
`wait_idle` with no other work, a multi-leg `send_all` can stall on a quiet
ring. Keep calling `submit()` from the issuer — tealetio already flushes
before `wait_idle`. Watching `ring.fd` does not see an unsubmitted SQE.
A worker that may not enter cannot submit the next-leg itself.
`DEFER_TASKRUN` already rejects worker `serve_completions`, so this pairing
is uncommon.
`IORING_SETUP_DEFER_TASKRUN` requires that same owning thread
to reap completions too: `wait()` and `serve_completions()` must run there,
not on a worker pool. Kernels expect `IORING_SETUP_DEFER_TASKRUN` together
with `IORING_SETUP_SINGLE_ISSUER`.

`IORING_SETUP_SQPOLL` enables a kernel submission-queue poller. Pass it in
`Ring(..., flags=...)` when you want that mode; ring construction may raise
`OSError` if the environment rejects it (privileges, container policy). There
is no dedicated capability key — handle failure at create time (or try
`probe(flags=IORING_SETUP_SQPOLL)` first if you prefer). Liburing's submit path
wakes a sleeping poller automatically when needed. When the SQ is full and the
poller has not yet freed a slot, prepare waits with the GIL released but still
under the ring critical section (up to a few seconds); prefer
`IORING_SETUP_SINGLE_ISSUER` (or a single submitter) with SQPOLL.

The compiled liburing version fields report the header version used to build the
binary extension. This is useful in CI because Linux distribution images can
compile the same Python package against different liburing development packages
while still running on the hosted runner's kernel.

`prepare_send_zc()` and `prepare_sendmsg_zc()` are best gated with
`probe()["IORING_OP_SEND_ZC"]` and `probe()["IORING_OP_SENDMSG_ZC"]`. Both
entries use the documented kernel 6.0 floor via `uname(2)`. Zerocopy can still
complete with `ENOTSUP` or `EOPNOTSUPP` for some protocols (for example
`AF_UNIX` on WSL); higher layers should route those sockets through copying
send paths. If your CI image is expected to support zerocopy on inet sockets,
make that expectation explicit:

```bash
uv run --active python - <<'PY'
import uring_api

probe = uring_api.probe()
print(probe)
raise SystemExit(0 if probe.get("IORING_OP_SEND_ZC") and probe.get("IORING_OP_SENDMSG_ZC") else 1)
PY
```

If the native extension cannot be imported after installation, importing
`uring_api` still succeeds and `probe()` returns `{}`. Source builds with
unsupported native dependencies warn and install the pure Python wrapper without
`_uring_api`.

The `IORING_ACCEPT_MULTISHOT` capability uses a runtime operation probe rather
than a kernel version check. It creates a private temporary ring and loopback
listener, submits one multishot accept request, connects a local client, and
checks whether the first accept completion keeps the request armed. If the build
headers do not expose the helper flag, the capability simply reports `False`.

The `IORING_POLL_MULTISHOT` capability uses a runtime operation probe. It creates
a private socket pair, submits one multishot poll for `POLLIN`, writes one byte
to the peer, and reports `True` only if the first completion reports readiness
and keeps the request armed with `IORING_CQE_F_MORE`. Gate
`prepare_poll_multishot()` on this entry; one-shot `prepare_poll()` and
`prepare_poll_remove()` are treated as baseline poll surface.

The `IORING_RECV_MULTISHOT` capability is also checked with a runtime operation
probe because it requires newer kernel support than multishot accept. It creates
a private socket pair and provided-buffer ring, submits one multishot receive,
sends one byte, and reports `True` only if the first completion selects a buffer
and keeps the request armed with `IORING_CQE_F_MORE`.

## Initialising a Ring

The current wrapper exposes the native ring lifecycle. A ring is a file
descriptor plus shared submission/completion queues owned by the process.

```python
import uring_api

with uring_api.Ring(entries=8) as ring:
    print("fd:", ring.fd)
    print("kernel features:", ring.features)
    print("submission entries:", ring.sq_entries)
    print("completion entries:", ring.cq_entries)
```

`entries` is the requested submission queue depth. The kernel may round or size
the actual submission and completion queues, so inspect `sq_entries` and
`cq_entries` after initialisation if the exact capacity matters.

Pass `flags=` to request setup modes that were accepted by `probe(flags=...)`:

```python
import uring_api

flags = uring_api.IORING_SETUP_SINGLE_ISSUER

if uring_api.probe(flags=flags):
    with uring_api.Ring(entries=8, flags=flags) as ring:
        ...
```

The constructor passes these flags to `io_uring_queue_init_params()` for the
real ring. The application is still responsible for the contracts implied by
each flag; for example, `IORING_SETUP_SINGLE_ISSUER` requires all submissions to
come from the owning thread.

If initialisation fails, the constructor raises `OSError`:

```python
import errno
import uring_api

try:
    ring = uring_api.Ring(entries=256)
except OSError as exc:
    if exc.errno == errno.EPERM:
        raise RuntimeError("io_uring is blocked by seccomp or policy") from exc
    if exc.errno == errno.ENOMEM:
        raise RuntimeError("io_uring could not allocate or pin the requested resources") from exc
    raise
else:
    try:
        print(ring.fd)
    finally:
        ring.close()
```

## Threading Model

`Ring` deliberately stays close to liburing's shared-ring model, but the Python
object adds native locking around the parts that matter for normal use.

The intended baseline is simple:

- one thread may reap completions with `wait()`;
- `poll()` is the same park as `wait()` without harvesting and without
    submitting: same thread rules and unique-waiter slot, but prepared SQEs
    stay queued until `submit()` or `wait()`, and it leaves the CQE for a
    later `wait()`.
    `break_wait()` unblocks `poll()` the same way it unblocks `wait()` (internal
    NOP; the following `wait()` may then return empty);
- other threads may call `construct_*` / `prepare_*`, `create_buf_group()`,
    and `break_wait()`. `poll()` follows the same caller-thread rules as `wait()`
    (any thread except `IORING_SETUP_DEFER_TASKRUN`, which is owner-only);
- `break_wait()` is safe to call while another thread is blocked in `wait()`
    or `poll()`;
- multiple concurrent `wait()` calls are serialised by the `Ring` object;
- alternatively, callers may start their own Python threads and have each one
    call `serve_completions()` to wait for completions and call the callback
    directly.

Rings created with `IORING_SETUP_DEFER_TASKRUN` do not follow that worker-pool
model. Submit, `wait()`, `poll()`, `serve_completions()`, and `break_wait()`
must all run on the owning thread established by the first gated call.

`break_wait()` is the single ring wakeup entry point. It always opens the
host-side `wait_idle()` park **immediately**. When completion service is not
active, it also best-effort submits **one** internal NOP (not a user completion)
so a caller blocked in `wait()` on an **empty CQ** can return. While
`serve_completions()` workers own CQ reaping, the NOP is skipped — only the idle
park is needed. `stop_serving()` still forces a NOP so workers blocked in the
kernel wait can observe stop.

`wait_idle` is a **multi-signaller, single-waiter** park: many threads may call
`break_wait()`, but only one host may park at a time (the proactor driver).
Concurrent `wait_idle` waiters are not supported.

If the submission queue is full, the NOP may be omitted and `break_wait()` still
succeeds: a full SQ means outstanding work, so a real CQE will arrive soon
enough. The idle park does not wait on the NOP path.

The NOP is not a broadcast to worker threads. Serve workers that drain a wake CQE
treat it as an empty/internal batch and continue.

Serving workers use the same receive side as `wait()`, so public `wait()` calls
raise `RuntimeError` while they are running. Each worker calls
`serve_completions()`, then loops until `stop_serving()` asks the service to
exit. One thread is the unique kernel waiter: it waits, consumes each ready CQE,
and either packs it or pushes a copy onto a work FIFO so other threads never
enter the completion queue. Each packer takes **one** CQE, drops the mutex, and
runs it to completion (package, `Ring.callback`, and any follow-up SQE that
fits). Packers do not `io_uring_submit` after a CQE (that unbatches the SQ).
The unique waiter always `io_uring_submit`s prepared SQEs before harvest when
this thread may enter, so next-leg prepares from delivery are entered without a
host `submit()`. A host `prepare` while the waiter is already in `wait_cqe`
still needs `submit()` (or `wait()`) to become kernel-visible. A send-all
next-leg is filled on the packer; that thread submits when it may enter and a
unique waiter may already be in `wait_cqe` (deadlock avoidance). Otherwise
harvest flush or host `submit()` publishes
it, or it parks on fill-wait if there is no SQ slot. Under `SINGLE_ISSUER` the issuer must keep flushing
(see the setup-flags caveat). `wait()` does the same consume-one path on the
calling thread and still returns the ready list (or delivers via
`Ring.callback`). Inline ``wait()`` with a callback flushes after the drain.
`stop_serving()` sets the
stop flag, wakes queue waiters, and uses `break_wait()` so a worker blocked
in the kernel wait can observe stop and exit. The caller owns the threads, so the
caller must join them before closing the ring; `close()` and `__exit__()` raise
while completion service is still active. `reset_serving()` clears the stop flag
so a fresh set of workers can enter `serve_completions()` again. If a delivery
callback raises, the ring invokes `exception_handler` when one is set. The handler
receives a context dict with `message`, `exception`, `ring`, and `completion`
(the CQE being delivered). When the handler returns normally, that worker
continues serving. When no handler is set, or the handler itself raises,
`serve_completions()` exits with the exception; only that worker stops — other
serving workers keep running until `stop_serving()`.

Native C clients can register a worker-thread callback through the C API. When a
C callback is present, the serving worker calls it instead of `Ring.callback`;
otherwise it falls back to the Python callback property.

```python
import uring_api
import threading


def delivered(completion):
    print(completion.user_data, completion.res, completion.result)


with uring_api.Ring() as ring:
    ring.callback = delivered
    threads = [threading.Thread(target=ring.serve_completions) for _ in range(2)]
    for thread in threads:
        thread.start()
    try:
        ring.prepare_recv(fd, bytearray(4096), 0, 200)
        ring.submit()
    finally:
        ring.stop_serving()
        for thread in threads:
            thread.join()
```

`close()` is still an owner-coordinated shutdown operation for submissions. Do
not close a ring while another thread may submit new user operations.

## C API

Native clients can include `uring_api_capi.h` and import `_uring_api._C_API` with
`PyCapsule_Import()`. Use `uring_api.get_include()` to find the installed header
directory when compiling an extension module.

The capsule currently exposes:

- `abi_version`, `struct_size`, and `feature_flags` for compatibility checks.
  While the package remains pre-release, `abi_version` stays at **1** but the
  function table may be reordered or extended; clients should compare
  `struct_size` and null-check pointers they rely on. **Break vs earlier v1
  drafts:** `ring_set_pre_submit` / `ring_set_c_pre_submit` were removed;
  all `ring_submit_*` / `ring_submit_*_nowait` op slots were dropped — C
  clients construct then `ring_prepare()`; `ring_construct_*_multishot` does
  not take `base_sequence` (set `completion.sequence` after construct). Python
  construct/prepare accept optional `base_sequence` after `user_data`.
  C completion callbacks receive one `Completion` per call (not a list).
  Appended: `completion_set_sequence`, `ring_wait_idle`,
  `completion_take_user_data`, `ring_poll`, `ring_stats`,
  `completion_arm_link_timeout`. `completion_clear_user_data` was removed
  (`take` covers it). Python `Ring.prepare_*` is construct+prepare sugar
  with cargo then `user_data`. Rebuild any out-of-tree C client that cached
  `offsetof` values;
- `compiled_liburing_major` and `compiled_liburing_minor` for build-time header
    visibility;
- `probe(entries, flags)`, which returns a new reference to the same flat
    availability and capability dictionary as `_uring_api.probe()`;
- `ring_new()`, lifecycle helpers, metadata helpers, `ring_construct_*()` for
    every waitable op, `statx_st_size()`, `ring_prepare()`,
    `completion_prepared()`, `completion_arm_link_timeout()` (any completion,
    before `ring_prepare`; `prepare` links the timeout SQE; `send_all`
    reapplies that same relative timeout on each leg, not as a drain deadline;
    `UringApiTimespec` is `tv_sec` plus `tv_nsec` in `0..999999999`, and
    `{0, 0}` is already expired; Python `Completion.timeout` stores the same
    value, in seconds),
    `completion_skip_success()`, `completion_set_skip_success()`,
    `completion_skip_all()`, `completion_set_skip_all()`,
    `ring_break_wait()`, `ring_wait()`, and `ring_poll()` (CQ-ready, no harvest);
- **not yet:** `BufGroup` lifecycle over the C API (`create_buf_group`,
    `close` / `release_callback`, C release hook). Provided-buffer constructs take
    a Python `BufGroup` object; manage groups from Python until that surface is
    added (see `ROADMAP.md`);
- `ring_set_callback()`, `ring_set_exception_handler()`, `ring_set_c_callback()`,
    `ring_serve_completions()`, `ring_stop_serving()`, and `ring_reset_serving()`
    for completion-service control;
- `completion_check()`, `completion_user_data()`, `completion_set_user_data()`,
    `completion_take_user_data()`,
    `completion_res()`, `completion_flags()`, `completion_sequence()`,
    `completion_set_sequence()`, `completion_result()`, and
    `completion_kind()` for native completion inspection. Kind values match
    `URING_API_COMPLETION_KIND_*` in `uring_api_completion_kinds.h` and
    `CompletionKind` in Python; `ring_wait_idle()` parks until `break_wait`;
- `ring_set_nowait_error_handler()` and `ring_submit()` (flush prepared SQEs).
    Tagged nowait is `completion_set_skip_all` then `ring_prepare`; error-only
    delivery is `completion_set_skip_success` then `ring_prepare` (no dedicated C
    nowait slots). `ring_auto_submit` / `ring_set_auto_submit` match `Ring.auto_submit`
    (default on; off raises `SubmissionQueueFull` instead of flushing a full SQ,
    and `wait()` does not auto-submit). The unique CQ waiter always submits
    before harvest when it may enter; TAKE never submits. There is no C
    accessor for that policy.
    `ring_stats()` fills `UringApiRingStats` with the same counters as `Ring.stats()`.

Check `URING_API_CAPI_FEATURE_CORE` before calling the function table. The flag
describes the capsule API surface, not runtime kernel support for individual
operations. Use `probe()` to check whether this process can create a ring and to
read runtime support for optional operation helpers from the returned flat
dictionary. A C completion callback receives the ring object, one user-visible
`Completion` per call, and the supplied `user_data`. Return `0` for success;
return a negative value with a Python exception set so the current
`serve_completions()` call exits with that error (other workers are not
stopped). Callback pointers must
not be changed while `serve_completions()` workers are active.

## Choosing Ring Sizes

Ring sizing is about queue depth, not payload buffer size. A modest application
can start with a small number of in-flight operations; a server usually wants
enough entries to cover its expected concurrent I/O without constantly draining
and refilling the ring.

Typical starting points:

| Use case | Suggested entries | Notes |
| --- | ---: | --- |
| Availability probe | 2 | Enough to prove the kernel will create a ring. |
| Modest local I/O | 8-32 | Good for simple tools and initial experiments. |
| Concurrent client work | 64-256 | Enough room for batches without large memory pressure. |
| Server-style I/O | 512-4096 | Needs deliberate resource-limit checks and backpressure. |

`UringProactor` still defaults to `entries=8` (modest / test size). Production
rings should pick from this table and may want `IORING_SETUP_CQSIZE` so a
multishot burst does not overflow the CQ. Create-time depth, `CQSIZE`,
`COOP_TASKRUN`, and recv/send hints (`POLL_FIRST`, `SOCK_NONEMPTY`) are
tracked in [ROADMAP.md](ROADMAP.md) under **Setup flags and SQ/CQ sizing**.

Ring entries and provided-buffer pools should be configured separately:

- ring entries control how many operations can be submitted or completed at
  once;
- `create_buf_group()` registers a provided-buffer ring whose storage stays
  pinned for receive operations that select buffers from that group;
- large provided-buffer pools can exceed `RLIMIT_MEMLOCK` even when ring
  creation itself succeeds.

`uring-api` does not yet expose fixed-buffer registration for send-side fixed
zero-copy variants. When that is added, treat it as a separate pool from
caller-owned `BufGroup` rings.

That distinction matters. During probing, a 64 MiB fixed-buffer pool exceeded a
default 64 MiB memlock limit because the limit must cover the pinned payload
memory plus kernel/accounting overhead.

You can inspect the process limit before choosing `BufGroup` sizes:

```python
import resource

soft, hard = resource.getrlimit(resource.RLIMIT_MEMLOCK)

print("memlock soft limit:", soft)
print("memlock hard limit:", hard)
```

Size provided-buffer pools explicitly rather than assuming the largest useful
value is safe:

```python
buffer_size = 16 * 1024
buffer_count = 256
pool_bytes = buffer_size * buffer_count

print("planned pinned buffer pool:", pool_bytes)
```

Good default `create_buf_group()` profiles would look something like:

| Profile | Ring entries | Buffer size | Buffer count | Pinned bytes |
| --- | ---: | ---: | ---: | ---: |
| modest | 32 | 16 KiB | 64 | 1 MiB |
| interactive | 128 | 16 KiB | 256 | 4 MiB |
| server | 1024 | 64 KiB | 1024 | 64 MiB |

The server profile is intentionally near the common default memlock limit on
some systems. In practice, leave headroom or raise the limit before registering
that much memory.

## Containers and Limits

Containers may block `io_uring_setup()` even when the host kernel supports it.
For example, Docker's default seccomp profile commonly rejects ring creation
with `EPERM`. A less restricted profile may be required for development.

Large `BufGroup` pools may also require raising `RLIMIT_MEMLOCK`. Prefer smaller
buffers while developing the operation model, then make server profiles opt-in
and explicit.

## Build Requirements

`uring-api` links against system `liburing`:

```bash
sudo apt install liburing-dev
```

The native extension requires `liburing >= 2.4`. Older headers do not expose the
version macros we use for build-time validation, and they also predate the data
and ring entry helpers used by the extension. On Ubuntu, that means
`ubuntu-23.10` or newer from distro packages; `ubuntu-22.04` needs a newer
liburing installed from another source to build `_uring_api`.

The extension uses multi-phase module initialisation and declares itself safe to
import without enabling the GIL on free-threaded CPython builds.
