Metadata-Version: 2.4
Name: autofission
Version: 0.0.3
Summary: Capacity-aware autoscaling limits for Fission on Kubernetes
Author: pomponchik
License: MIT
Project-URL: Documentation, https://github.com/pomponchik/autofission
Project-URL: Source, https://github.com/pomponchik/autofission
Project-URL: Tracker, https://github.com/pomponchik/autofission/issues
Keywords: fission,kubernetes,autoscaling,serverless
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: System Administrators
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: 3.15
Classifier: Programming Language :: Python :: Free Threading
Classifier: Programming Language :: Python :: Free Threading :: 3 - Stable
Classifier: Topic :: System :: Clustering
Classifier: Typing :: Typed
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: kubernetes<36.0.0,>=27.2.0; python_version < "3.10"
Requires-Dist: kubernetes<37.0.0,>=36.0.3; python_version >= "3.10"
Requires-Dist: urllib3<3.0.0,>=1.26.0
Dynamic: license-file

<details>
  <summary>ⓘ</summary>

[![Downloads](https://static.pepy.tech/badge/autofission/month)](https://pepy.tech/project/autofission)
[![Downloads](https://static.pepy.tech/badge/autofission)](https://pepy.tech/project/autofission)
[![Coverage Status](https://coveralls.io/repos/github/pomponchik/autofission/badge.svg?branch=develop)](https://coveralls.io/github/pomponchik/autofission?branch=develop)
[![Lines of code](https://sloc.xyz/github/pomponchik/autofission/?category=code)](https://github.com/boyter/scc/)
[![Hits-of-Code](https://hitsofcode.com/github/pomponchik/autofission?branch=develop)](https://hitsofcode.com/github/pomponchik/autofission/view?branch=develop)
[![Test-Package](https://github.com/pomponchik/autofission/actions/workflows/tests_and_coverage.yml/badge.svg?branch=develop)](https://github.com/pomponchik/autofission/actions/workflows/tests_and_coverage.yml)
[![Python versions](https://img.shields.io/pypi/pyversions/autofission.svg)](https://pypi.org/project/autofission/)
[![PyPI version](https://badge.fury.io/py/autofission.svg)](https://pypi.org/project/autofission/)
[![Checked with mypy](http://www.mypy-lang.org/static/mypy_badge.svg)](http://mypy-lang.org/)
[![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff)
[![DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/pomponchik/autofission)

</details>

![Autofission](https://raw.githubusercontent.com/pomponchik/autofission/develop/docs/assets/logo.svg)


Autofission dynamically calculates and updates the maximum replica limit (`MaxScale`) for explicitly opted-in [Fission](https://fission.io/) Functions. It estimates the limit from currently available CPU, memory, and Pod slots on schedulable Kubernetes nodes.

A Fission Function using the [`newdeploy` executor](https://fission.io/docs/usage/function/executor/) can scale down when demand disappears. Its [Horizontal Pod Autoscaler (HPA)](https://kubernetes.io/docs/concepts/workloads/autoscaling/) still needs a fixed positive `MaxScale`. A limit sized for today's cluster becomes too low when nodes are added. An arbitrarily high limit can flood the scheduler with Pods that cannot fit.

Autofission keeps the limit current by estimating available capacity on each node. It subtracts the requests of existing workloads and accounts for the Function's current replicas. It is designed for elastic, bare-metal, homelab, and edge clusters where nodes come and go and idle compute should remain available to Functions without scheduler preemption of existing services.

The only scaling setting Autofission changes is this upper bound; it also records the calculation in annotations. Fission's HPA and idle reaper still decide when the replica count grows and shrinks.

## Table of Contents

- [**Installation**](#installation)
- [**Quick start**](#quick-start)
- [**How it works**](#how-it-works)
- [**Configuration**](#configuration)
- [**RBAC and security**](#rbac-and-security)
- [**Operations**](#operations)
- [**Compatibility and limitations**](#compatibility-and-limitations)
- [**Troubleshooting**](#troubleshooting)


## Installation

Use the Python CLI for optional local or one-shot runs. Install the Helm chart to run Autofission continuously in a cluster.

### Python CLI

For local use, Autofission requires Python 3.8 or newer. CI tests Python 3.8 through 3.14, free-threaded Python 3.14, and Python 3.15 beta:

```bash
python -m pip install autofission
autofission --help
```

With no credential-selection flags, the CLI uses the current `kubeconfig` context unless `KUBERNETES_SERVICE_HOST` is present, in which case it uses mounted ServiceAccount credentials. `--in-cluster` selects the ServiceAccount explicitly.

To run one reconciliation without starting a daemon:

```bash
autofission --once --context my-cluster
```

This performs a real reconciliation and can update opted-in Functions.

### Install in a cluster

Prerequisites are Kubernetes and an [existing Fission installation](https://fission.io/docs/installation/). Autofission manages only Functions that use the `newdeploy` executor. By default, the chart installs the controller, its RBAC, and two PriorityClasses; it does not install or remove Fission.

Set `AUTOFISSION_VERSION` to the chart version you want to install, then install Autofission from its OCI release:

```bash
helm upgrade --install autofission \
  oci://ghcr.io/pomponchik/charts/autofission \
  --version "${AUTOFISSION_VERSION}" \
  --namespace fission \
  --create-namespace \
  --atomic \
  --wait
```

For an unreleased checkout, build a controller image that the cluster can pull, replace the OCI URL with `./deploy/helm/autofission`, and set its repository and tag. Autofission assumes fetcher requests of `10m` CPU and `16Mi` memory. If Fission uses different values, pass them to the chart so the capacity calculation remains accurate:

```bash
helm upgrade --install autofission ./deploy/helm/autofission \
  --namespace fission \
  --set-string image.repository=REGISTRY/autofission \
  --set-string image.tag=TAG \
  --set controller.fetcherCpuRequest=20m \
  --set controller.fetcherMemoryRequest=32Mi
```

After the chart creates its PriorityClasses, configure Fission to use the [low, non-preempting runtime class](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/). Its default value of `-10` places unprioritized Pods, whose priority is normally `0`, ahead of elastic Function Pods in the scheduling queue. With the default Autofission release name, add the following to Fission's Helm values and upgrade Fission before opting in any Functions:

```yaml
runtimePodSpec:
  enabled: true
  podSpec:
    priorityClassName: autofission-runtime
```


## Quick start

The example uses a Function named `hello` in the `default` namespace and assumes Fission's default same-namespace workload placement. Replace the Function name and namespace, then opt it in. If Fission sets a separate `functionNamespace`, use that workload namespace when querying the HPA:

```bash
kubectl label function hello \
  --namespace default \
  autoscaling.fission.io/cluster-capacity=true
```

After the next successful reconciliation, inspect the Function's `MaxScale`, the recorded calculation, and the HPA. The daemon waits 15 seconds between cycles by default; processing and transient failures can add delay:

```bash
kubectl get function hello --namespace default \
  -o jsonpath='{.spec.InvokeStrategy.ExecutionStrategy.MaxScale}{"\n"}'
kubectl get function hello --namespace default \
  -o jsonpath='{.metadata.annotations.autoscaling\.fission\.io/calculated-maxscale}{"\n"}'
kubectl get hpa --namespace default
```

After a successful cycle, the first two values should match. The HPA may reflect the new limit slightly later because Fission updates it asynchronously.

Removing the label stops future management. Autofission deliberately does not guess or restore a previous manually configured `MaxScale`; the last calculated value and informational annotations remain until you change them.


## How it works

```mermaid
flowchart TD
    state["Cluster state<br/>Nodes, Pods, and resource requests"] --> autofission["Autofission<br/>estimates modeled capacity"]
    autofission -->|updates| limit["Function MaxScale"]
    demand["Demand or idle time"] --> scaling["Fission HPA<br/>and idle reaper"]
    limit -->|sets upper bound| scaling
    scaling -->|changes| replicas["Function replicas"]
```

Each cycle validates its inputs and is idempotent for an unchanged cluster snapshot:

1. List opted-in Functions, along with Fission Environments, Kubernetes Nodes, and Pods.
2. Keep `Ready`, uncordoned, non-deleting nodes. Nodes with `NoSchedule` or `NoExecute` taints are excluded by default.
3. Calculate effective [CPU and memory requests](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/) for every active, scheduled Pod, including containers, init containers and sidecars, Pod-level requests, and overhead. Exclude completed and unbound Pods.
4. Resolve each Function's CPU and memory requests, inheriting missing or zero values from its Environment, then add the fetcher request. Use larger values observed on an existing Function Pod.
5. Add the Function's existing Pods back to its capacity budget because step 3 counted them as other workload. This changes only the calculation, not the Pods. Then calculate how many identical Pods fit on each node. Per-node results are summed, so CPU on one node cannot combine with memory on another.
6. Set `MaxScale` to at least `MinScale` and `1`. Patch a Function when either `MaxScale` or Autofission's calculation annotations have changed. A `resourceVersion` conflict prevents concurrent edits from being overwritten.

An invalid or conflicting Function does not block the others. The controller becomes `Ready` after a cycle in which every managed Function completes cleanly. A later failure does not clear readiness immediately: the last successful marker remains valid until the configured probe age expires (`60` seconds by default).


## Configuration

CLI flags take precedence over non-empty environment variables, which take precedence over defaults.

| CLI flag | Environment variable | Helm value | Default |
|---|---|---|---|
| `--interval-seconds` | `AUTOFISSION_INTERVAL_SECONDS` | `controller.intervalSeconds` | `15` |
| `--request-timeout-seconds` | `AUTOFISSION_REQUEST_TIMEOUT_SECONDS` | `controller.requestTimeoutSeconds` | `10` |
| `--retry-attempts` | `AUTOFISSION_RETRY_ATTEMPTS` | `controller.retryAttempts` | `3` |
| `--fetcher-cpu-request` | `AUTOFISSION_FETCHER_CPU_REQUEST` | `controller.fetcherCpuRequest` | `10m` |
| `--fetcher-memory-request` | `AUTOFISSION_FETCHER_MEMORY_REQUEST` | `controller.fetcherMemoryRequest` | `16Mi` |
| `--managed-label` | `AUTOFISSION_MANAGED_LABEL` | `controller.managedLabel` | `autoscaling.fission.io/cluster-capacity` |
| `--managed-value` | `AUTOFISSION_MANAGED_VALUE` | `controller.managedValue` | `true` |
| `--include-tainted-nodes` | `AUTOFISSION_INCLUDE_TAINTED_NODES` | `controller.includeTaintedNodes` | `false` |
| `--log-level` | `AUTOFISSION_LOG_LEVEL` | `controller.logLevel` | `INFO` |
| `--state-directory` | `AUTOFISSION_STATE_DIRECTORY` | — | `/tmp/autofission` (CLI); `/var/run/autofission` (chart) |

`--kubeconfig`, `--context`, and `--in-cluster` select credentials. `--once` runs exactly one pass. `--probe readiness|liveness --max-age-seconds N` is intended for Kubernetes exec probes.


## RBAC and security

With the default RBAC values, the chart-created ClusterRole grants only these cluster-wide operations:

- `list` Nodes and Pods;
- `list` Fission Environments;
- `list` and `patch` Fission Functions.

The chart does not grant permission to read Secrets, create or delete Functions, or mutate Pods, Nodes, Deployments, or Services. Other bindings attached to the same ServiceAccount can grant additional permissions. Kubernetes RBAC cannot restrict `list` or `patch` permissions by label, so the opt-in label is an application-level boundary. Listing Pods exposes their specifications, including literal environment-variable values, to the controller process. This access is needed to protect capacity requested by other workloads.

By default, the container runs as UID/GID `65532` with a read-only root filesystem. It drops all Linux capabilities, blocks privilege escalation, and uses a `RuntimeDefault` seccomp profile. The chart creates a NetworkPolicy with an empty ingress list and no egress policy. Enforcement requires a compatible network plugin, and other NetworkPolicies can add allowed ingress because Kubernetes combines their rules.

The controller has a high, non-preempting PriorityClass, so a pending controller Pod is queued ahead of lower-priority Pods when capacity becomes available. It cannot evict running Pods, so the class does not guarantee availability after a node failure or in a full cluster. The `autofission-runtime` class is negative and also uses `preemptionPolicy: Never`; configure Fission to use it as shown under installation. Accurate resource requests remain essential because Autofission budgets requests like the scheduler rather than measuring live CPU or memory usage.

Treat permission to set the opt-in label as permission to consume the cluster's elastic budget. In a multi-tenant cluster, restrict that label with your admission policy. Autofission does not implement cross-Function quotas or fair sharing.


## Operations

The chart runs one replica with a `Recreate` strategy, preventing rollout overlap within one installation without requiring leader election. After a node or cluster restart, the Deployment restores its controller replica and Autofission rebuilds its state from the API.

Useful checks for the default release name, namespace, and ServiceAccount:

```bash
kubectl rollout status deployment/autofission --namespace fission
kubectl auth can-i list pods \
  --as=system:serviceaccount:fission:autofission --all-namespaces
kubectl auth can-i patch functions.fission.io \
  --as=system:serviceaccount:fission:autofission --all-namespaces
kubectl auth can-i get secrets \
  --as=system:serviceaccount:fission:autofission --all-namespaces
```

With the default chart-managed RBAC and no additional bindings, the rollout should complete; the Pod-list and Function-patch checks should say `yes`, and the Secret check should say `no`.

Before uninstalling, remove the opt-in labels and set each Function's `MaxScale` to the limit you want to retain: stopping Autofission does not restore older values.

Uninstalling removes the controller resources but retains the runtime PriorityClass because Fission may still reference it during future cold starts. Fission CRDs, Functions, Environments, and namespaces remain untouched. Remove the PriorityClass manually only after ensuring that Fission's global `runtimePodSpec` and all Function or Environment PodSpecs no longer reference it.


## Compatibility and limitations

- Only Fission `newdeploy` Functions are managed. `poolmgr`, empty executor, and `container` are rejected as opt-in configuration errors.
- Every managed Function receives the full capacity it could use by itself. This preserves burst capacity, but simultaneous cold bursts can temporarily create `Pending` Pods. Later cycles account for scheduled peers and reduce the limits; Autofission is not a fairness scheduler.
- Fission and Kubernetes require a positive HPA maximum, so capacity below one replica produces `MaxScale=1`. An explicit `MinScale` is honored even when it exceeds currently free capacity.
- The [Fission v1 API](https://fission.io/docs/reference/crd-reference/) is tested end to end with Fission `1.27.0`. Environment resource inheritance follows Fission's override semantics and is covered by unit tests.
- Capacity includes CPU, memory, and Pod slots. It does not model storage, GPUs and other extended resources, quotas, topology or affinity, per-Function scheduling constraints, image architecture, or unscheduled third-party Pods.
- Tainted nodes are excluded unless `--include-tainted-nodes` is explicitly set. Only enable it when Fission runtime Pods actually tolerate those taints.
- Extra sidecars and runtime `PodSpec` overhead are learned from a running Function Pod. Before the first replica, the estimate consists of the resolved runtime container plus configured fetcher requests.
- Low, non-preempting priority prevents Function Pods from evicting existing workloads. It cannot prevent node-pressure eviction when requests are inaccurate or nodes run at their physical limit; reserve headroom and set accurate requests.
- Very large clusters should benchmark API-server load and controller memory before shortening the default interval.
- Cluster-scoped RBAC and PriorityClass names use the release name, so reusing a name in another namespace causes collisions. Prefer one Autofission installation per cluster. Multiple installations require unique `fullnameOverride` and `priorityClasses.*.name` values plus selectors that assign disjoint sets of Functions to their `managedLabel` and `managedValue` pairs; two controllers must never manage the same Function.


## Troubleshooting

`Autofission is NotReady` — inspect controller logs. A `403` usually indicates missing custom RBAC or a ServiceAccount mismatch. A `409` means the Function changed after it was listed; the daemon will retry it on the next cycle. Other errors name the Function where possible.

`The calculated limit is smaller than expected` — check cordons, `Ready` status, taints, Pod requests, Pod slots, fetcher values, and per-node fragmentation. Capacity cannot combine spare CPU and spare memory located on different nodes.

`Function Pods remain Pending` — verify the runtime PriorityClass, taints/tolerations, node selectors, architecture, quota, storage, and concurrent managed Functions. Those constraints can be stricter than Autofission's current capacity model.

`The HPA has not changed yet` — Autofission patches the Function CR. Fission's executor reconciles that change into the HPA asynchronously; controller readiness confirms only that Autofission completed a full reconciliation successfully within the configured probe age.
