# Eigrel

> A declarative language for tabular data and machine learning: describe datasets,
> transformations, features and models in one small, fully-checked syntax, and the compiler
> generates pandas + scikit-learn, PySpark or SQL from it. Programs are static: every column
> reference, type and algorithm parameter is verified before anything runs, so Eigrel is a
> safe generation target for LLMs — write `.eig`, let `eigrel check` prove it, let the
> compiler decide how it runs.

## Agent workflow

Verify-before-run loop for generating Eigrel programs:

1. Write the program to a `.eig` file, following the language reference below.
2. Run `eigrel check --json FILE.eig`. Exit code 0 means the program is valid; exit code 1
   means every `files[].errors[]` entry is a real mistake to fix.
3. Each error has a `stage` (`lex`, `parse`, `semantic`, `io`), a `message`, and the exact
   `line` and `column` (1-based) it points at. Fix the first error and check again; later
   errors may be consequences of the first one.
4. Once `ok` is `true`, the program compiles for sure. Before running it, run
   `eigrel plan --json FILE.eig` (add `--target spark` for Spark). The plan reads the real data:
   it re-checks the program against the actual column names and types, counts the rows every
   step keeps, and returns `findings[]`, each with a `severity` (`error`, `warning`, `info`), a
   stable `code` (e.g. `schema`, `target-missing`, `missing-values`, `imbalance`,
   `small-validation`, `drift-removed`) and the `line` it concerns. Errors predict a failed run:
   fix them. Warnings are judgement calls to report to the human.
5. Then:
   - `eigrel compile FILE.eig` prints the generated Python (add `--target spark|sql`),
   - `eigrel run FILE.eig` runs it (needs `pip install "eigrel[python]"`),
   - `eigrel ir FILE.eig` shows the operation graph the backends work from.

Agents that speak the Model Context Protocol can run `eigrel mcp` (with
`pip install "eigrel[mcp]"`) and call the `check`, `plan`, `compile` and `run` tools instead of
the CLI; they return the same JSON. This file is also served as the `eigrel://llms.txt` resource.

Example error object:

```json
{
  "stage": "semantic",
  "message": "column 'income' does not exist here; available columns: age, purchases",
  "line": 9,
  "column": 5
}
```

The compiler is the source of truth: if `check` passes, generated code is deterministic and
never invents columns, parameters or algorithms.

## Language reference

A program is a list of statements, each starting with one of seven keywords:
`dataset`, `transform`, `features`, `model`, `train`, `evaluate`, `register`.
There are no variables, no loops and no function calls. Comments start with `#`.

### dataset

```eigrel
dataset customers from csv("data/customers.csv")
dataset events    from parquet("data/events.parquet")
dataset reads     from json("data/reads.jsonl")            # .json/.jsonl/.ndjson
dataset orders    from sql(env("DATABASE_URL"), "shop.orders")
dataset users     from bigquery("my-project.analytics.users")
```

`env("NAME")` reads a secret from the environment at run time (only as a SQL URL).

### transform

Mutates the named dataset in order; each operation produces a new version of it.

```eigrel
transform customers {
    fill income = 0, city = "unknown"    # literal values for missing values
    drop_missing age, purchases          # drop rows with a missing value in these columns
    drop_missing                          # ...or in any column
    filter age >= 18                      # keep rows matching a boolean expression
    select age, income, purchases, churned         # keep only these columns
}
```

Expressions in `filter`: column names, integer/float/string/boolean literals (`true`,
`false`, strings in double quotes with `\"` and `\\` escapes), parentheses, `not`,
comparisons `== != < <= > >=`, arithmetic `+ - * / %`, and `and` / `or`. Comparisons do not
chain (write `x > 1 and x < 9`); function calls are not allowed; the condition must be
boolean.

### features

The columns used to train. Optional; without it a model trains on every column except the
target.

```eigrel
features customers {
    age
    income
    purchases
}
```

### model

Declares a model from an algorithm. `task = classification|regression` is inferred from the
metrics used in `evaluate` when omitted.

```eigrel
model churn = random_forest {
    trees = 100
}
```

| algorithm | task(s) | parameters |
|---|---|---|
| random_forest | classification, regression | trees, max_depth, min_samples_leaf |
| decision_tree | classification, regression | max_depth, min_samples_leaf |
| gradient_boosting | classification, regression | trees, learning_rate, max_depth |
| xgboost | classification, regression | trees, learning_rate, max_depth |
| logistic_regression | classification | max_iter, c |
| linear_regression | regression | (none) |

All parameters are positive numbers; `trees`/`max_depth`/`min_samples_leaf`/`max_iter` are
integers.

### train

```eigrel
train churn {
    target = churned      # required
    data = customers      # optional; inferred when exactly one dataset fits
    validation = 0.2      # optional fraction (0..1); default 0.2
    seed = 42             # optional; default 42
}
```

The target column must not be listed in `features`. A model can be trained once.

### evaluate

```eigrel
evaluate churn {
    metrics = [accuracy, precision, recall, f1]   # optional; defaults below
}
```

| task | metrics |
|---|---|
| classification | accuracy, precision, recall, f1, auc |
| regression | mae, mse, rmse, r2 |

Defaults when `metrics` is omitted: `accuracy, precision, recall, f1` (classification) and
`mae, rmse, r2` (regression). Metrics must match the model's task. Evaluating before
`register` logs the metrics to the MLflow run.

### register

Registers the trained model in MLflow. The whole pipeline (preprocessing + estimator) is
registered, so the model version accepts the raw columns the program trained on.

```eigrel
register churn {
    name = "customer-churn"        # optional; default: the model name
    experiment = "eigrel-examples" # optional MLflow experiment
}
```

### assumptions

Declare a dataset's time column; every model trained on it then validates on the latest rows.

```eigrel
assumptions customers {
    time = signup_date
}
```

### predict

Score a dataset with a trained model. The data must provide every training feature with the same
types; training fills are replayed. Writes `<target>_prediction` (and `<target>_probability`).

```eigrel
predict churn {
    data = new_customers
    output = csv("scored.csv")
}
```

## Examples

Complete programs, each with the request it answers. They pass `eigrel check`.

Request: "Predict which customers churn from data/customers.csv. Use XGBoost, treat a missing
income as 0, report accuracy, f1 and AUC, and register the model as customer-churn."

```eigrel
dataset customers from csv("data/customers.csv")

transform customers {
    fill income = 0
    filter age >= 18
}

features customers {
    age
    income
    purchases
}

model churn = xgboost {
    trees = 200
    learning_rate = 0.05
}

train churn {
    target = churned
}

evaluate churn {
    metrics = [accuracy, f1, auc]
}

register churn {
    name = "customer-churn"
}
```

Request: "Forecast order value from the orders table in our Postgres database (the URL is in
DATABASE_URL). Validate on the most recent orders, not a random sample."

```eigrel
dataset orders from sql(env("DATABASE_URL"), "shop.orders")

transform orders {
    drop_missing order_value, ordered_at
    select ordered_at, items, discount, channel, order_value
}

assumptions orders {
    time = ordered_at
}

model value = gradient_boosting {
    trees = 300
    max_depth = 4
}

train value {
    target = order_value
    validation = 0.2
}

evaluate value {
    metrics = [mae, rmse, r2]
}
```

Request: "Train a churn model on last quarter's customers and score the new sign-ups in
data/new.csv, writing the predictions to data/scored.csv."

```eigrel
dataset customers from csv("data/customers.csv")

model churn = random_forest {
    trees = 200
}

train churn {
    target = churned
}

dataset signups from csv("data/new.csv")

predict churn {
    data = signups
    output = csv("data/scored.csv")
}
```

## Semantics the compiler enforces (fix these before checking)

- Every referenced column must exist. Columns are known only after a `select` or from a
  `features` block; reference only columns you have declared there.
- `fill` values must be literals; `fill`/`select`/`drop_missing` columns must exist.
- The `target` must not be a feature of the same dataset.
- Algorithm names, parameters and metrics are exact (see tables above).
- Keywords (`dataset transform features model train evaluate register fill drop_missing
  and or not true false`) cannot be used as names.

## Files and links

- [Language reference](https://github.com/thentsation/eigrel/blob/main/docs/LANGUAGE.md): the
  human-facing version of the syntax, including backend differences.
- [Examples](https://github.com/thentsation/eigrel/tree/main/examples): minimal programs,
  including `mlflow.eig` (fill/drop_missing/xgboost/register).
- [CLI](https://github.com/thentsation/eigrel#cli): `init`, `check --json`, `plan --json`, `run`,
  `compile`, `ir`, `ast`, `tokens`.
