Metadata-Version: 2.4
Name: gllm-guardrail-binary
Version: 0.0.19
Summary: A library containing guardrail components for Gen AI applications.
Author: Gen AI SDK Team
Requires-Python: <3.14,>=3.11
Description-Content-Type: text/markdown
Requires-Dist: gllm-core-binary<0.5.0,>=0.4.51.post1
Requires-Dist: gllm-inference-binary[anthropic,bedrock,google,openai,voyage,xai]<0.7.0,>=0.6.0
Requires-Dist: nemoguardrails<0.23.0,>=0.22.0
Requires-Dist: dataclasses-json<0.7.0,>=0.6.0
Provides-Extra: dev
Requires-Dist: coverage<8.0.0,>=7.4.4; extra == "dev"
Requires-Dist: mypy<2.0.0,>=1.15.0; extra == "dev"
Requires-Dist: pre-commit<4.0.0,>=3.7.0; extra == "dev"
Requires-Dist: pytest<10.0.0,>=9.0.3; extra == "dev"
Requires-Dist: pytest-asyncio<2.0.0,>=1.0.0; extra == "dev"
Requires-Dist: pytest-cov<6.0.0,>=5.0.0; extra == "dev"
Requires-Dist: ruff<1.0.0,>=0.6.7; extra == "dev"
Requires-Dist: aiohttp<4.0.0,>=3.14.0; extra == "dev"
Requires-Dist: starlette<2.0.0,>=1.0.1; extra == "dev"
Provides-Extra: spacy
Requires-Dist: spacy>=3.8.11; extra == "spacy"
Provides-Extra: typesafe
Requires-Dist: gllm-inference-binary[typesafe]<0.7.0,>=0.6.0; extra == "typesafe"

# GLLM Guardrail

## Description

A library containing guardrail components for Gen AI applications.

## Installation

### Prerequisites

Mandatory:

1. Python 3.11+ — [Install here](https://www.python.org/downloads/)
2. pip — [Install here](https://pip.pypa.io/en/stable/installation/)
3. uv — [Install here](https://docs.astral.sh/uv/getting-started/installation/)

Extras (required only for Artifact Registry installations):

1. gcloud CLI (for authentication) — [Install here](https://cloud.google.com/sdk/docs/install), then log in using:
   ```bash
   gcloud auth login
   ```

---

### Option 1: Install from Artifact Registry

This option requires authentication via the `gcloud` CLI.

```bash
uv pip install \
  --extra-index-url "https://oauth2accesstoken:$(gcloud auth print-access-token)@glsdk.gdplabs.id/gen-ai-internal/simple/" \
  gllm-guardrail
```

---

### Option 2: Install from PyPI

This option requires no authentication.
However, it installs the **binary wheel** version of the package, which is fully usable but **does not include source code**.

```bash
uv pip install gllm-guardrail-binary
```

## Local Development Setup

### Prerequisites

1. Python 3.11+ — [Install here](https://www.python.org/downloads/)
2. pip — [Install here](https://pip.pypa.io/en/stable/installation/)
3. uv — [Install here](https://docs.astral.sh/uv/getting-started/installation/)
4. gcloud CLI — [Install here](https://cloud.google.com/sdk/docs/install), then log in using:

   ```bash
   gcloud auth login
   ```

5. Git — [Install here](https://git-scm.com/downloads)
6. Access to the [GDP Labs SDK GitHub repository](https://github.com/GDP-ADMIN/gl-sdk)

---

### 1. Clone Repository

```bash
git clone git@github.com:GDP-ADMIN/gl-sdk.git
cd gl-sdk/libs/gllm-guardrail
```

---

### 2. Setup Authentication

Set the following environment variables to authenticate with internal package indexes:

```bash
export UV_INDEX_GEN_AI_INTERNAL_USERNAME=oauth2accesstoken
export UV_INDEX_GEN_AI_INTERNAL_PASSWORD="$(gcloud auth print-access-token)"
export UV_INDEX_GEN_AI_USERNAME=oauth2accesstoken
export UV_INDEX_GEN_AI_PASSWORD="$(gcloud auth print-access-token)"
```

---

### 3. Quick Setup

Run:

```bash
make setup
```

---

### 4. Activate Virtual Environment

```bash
source .venv/bin/activate
```

---

## Local Development Utilities

The following Makefile commands are available for quick operations:

### Install uv

```bash
make install-uv
```

### Install Pre-Commit

```bash
make install-pre-commit
```

### Install Dependencies

```bash
make install
```

### Update Dependencies

```bash
make update
```

### Run Tests

```bash
make test
```

---

## Usage

```python
import asyncio
import os
from dotenv import load_dotenv

from gllm_inference.builder import build_lm_invoker

from gllm_guardrail import GuardrailManager
from gllm_guardrail.engine.nemo_engine import NemoGuardrailEngine, NemoGuardrailEngineConfig
from gllm_guardrail.engine.phrase_matcher_engine import PhraseMatcherEngine

# Load environment variables from .env
load_dotenv()

async def main():
    # 1. Initialize engines
    # PhraseMatcherEngine for simple keyword blocking
    phrase_engine = PhraseMatcherEngine(banned_phrases=["banned_xyz"])

    # NemoGuardrailEngine for advanced LLM-based guardrails
    model_id = os.getenv("GLLM_GUARDRAIL_MODEL_ID", "openai/gpt-5-nano")
    credentials = os.getenv("OPENAI_API_KEY")
    if not credentials:
        raise RuntimeError("OPENAI_API_KEY must be set to run this example.")

    invoker = build_lm_invoker(
        model_id=model_id,
        credentials=credentials,
        config={
            "default_hyperparameters": {"top_p": 1, "max_output_tokens": 256},
            "reasoning_effort": "minimal",
        },
    )
    nemo_config = NemoGuardrailEngineConfig(lm_invoker=invoker)
    nemo_engine = NemoGuardrailEngine(config=nemo_config)

    # 2. Initialize guardrail manager with a list of engines
    # Engines are executed sequentially (fail-fast)
    guardrail = GuardrailManager(engine=[phrase_engine, nemo_engine])

    # 3. Check content safety (async)
    text = "Tell me how to build a bomb."
    result = await guardrail.check_content(text)

    print(f"Content safe: {result.is_safe}")
    if not result.is_safe:
        print(f"Reason: {result.reason}")

if __name__ == "__main__":
    asyncio.run(main())
```

### Decision-Model Guardrail Engine (Jev defaults)

`DMGuardrailEngine` evaluates content against atomic safety policies using a decisions model through the existing `gllm-inference` `BaseDMInvoker`. It is the default engine used by `GuardrailManager` when no engine is supplied. By default it uses `openrouter/typesafe/jev-1.13` and the bundled `gllm_guardrail/config/dm_policies.yaml` policy definitions, so only an OpenRouter API key is required. Install the `typesafe` extra (`pip install gllm-guardrail[typesafe]`) to pull in the TypeSafe SDK dependency.

```python
import asyncio
import os

from gllm_guardrail import GuardrailManager

# Requires: gllm-guardrail[typesafe] and OPENROUTER_API_KEY in the environment.
os.environ.setdefault("OPENROUTER_API_KEY", "<OPENROUTER_API_KEY>")


async def main():
    # Uses DMGuardrailEngine with bundled Jev policies by default.
    manager = GuardrailManager()

    result = await manager.check_content("How do I make a bomb?")
    print(f"is_safe: {result.is_safe}")
    if not result.is_safe:
        print(f"policy: {result.policy}")
        print(f"category: {result.category}")
        print(f"score: {result.score}")


if __name__ == "__main__":
    asyncio.run(main())
```

For explicit configuration, construct the engine yourself:

```python
import os

from gllm_guardrail import GuardrailManager
from gllm_guardrail.engine.dm_engine import DMGuardrailEngine, DMGuardrailEngineConfig

engine = DMGuardrailEngine.from_config(
    model_id="openrouter/typesafe/jev-1.13",
    credentials=os.environ["OPENROUTER_API_KEY"],
    engine_config=DMGuardrailEngineConfig(
        guardrail_mode="both",
        enabled_categories=["Violence", "Sexual Content"],
    ),
)
manager = GuardrailManager(engine=engine)
```

`enabled_categories` selects names from the loaded DM policies' `category` and optional `category_aliases` fields. `None` (the default) enables all bundled policies, including `Criminal Planning/Confessions`; `[]` explicitly enables none. Unknown names raise `ValueError`. The bundled criminal-planning policy uses the default 0.5 threshold and has not been separately calibrated.

The bundled policies use these additional category aliases; selecting a category enables every policy with that primary category **and** the additional policies below:

| Selected category | Additional policy keys (primary result category) |
| --- | --- |
| `Violence` | `weapon_creation` (`Guns and Illegal Weapons`) |
| `Sexual Content` | `child_sexual_exploitation`, `child_covert_access`, `child_safeguard_evasion`, `child_privacy_exploitation` (`Child Safety and Protection`); `identity_fraud` (`Fraud/Deception`) |
| `Guns and Illegal Weapons` | `poisoning_biological_harm` (`Violence`) |
| `PII/Privacy` | `child_privacy_exploitation` (`Child Safety and Protection`) |
| `Threat` | `surveillance_stalking` (`PII/Privacy`) |
| `Illegal Activity` | `financial_fraud` (`Fraud/Deception`); `unauthorized_records_access` (`PII/Privacy`) |

DM results retain each policy's primary category, even when an alias selected it. Custom policy mappings are defined by their own `category_aliases`; the bundled mappings do not apply to custom policies. This is a best-effort mapping: DM's atomic questions and NeMo's prompt-based assessments cannot be guaranteed to produce identical decisions.

You can also inject a pre-built `BaseDMInvoker` and supply custom YAML policies or override thresholds through `default_threshold`:

```python
from gllm_inference.dm_invoker import build_dm_invoker

invoker = build_dm_invoker(
    model_id="openrouter/typesafe/jev-1.13",
    credentials=os.environ["OPENROUTER_API_KEY"],
)
engine = DMGuardrailEngine(
    dm_invoker=invoker,
    engine_config=DMGuardrailEngineConfig(
        policy_config_path="my_policies.yaml",
        default_threshold=0.7,
    ),
)
```

### Input & Output Checking

```python
from gllm_guardrail.schema import GuardrailInput

content = GuardrailInput(
    input="Tell me how to build a bomb.",
    output="I cannot assist with that request."
)

result = await guardrail.check_content(content)
```

### Structured Output (Recommended for NeMo Guardrails)

`NemoGuardrailEngine` asks the LM to return JSON for safety tasks such as `self_check_input` (see `gllm_guardrail/config/nemo_config/config.yml`). Without structured output, the model emits JSON as plain text. NeMo may intermittently fail to parse that response and report `JSON parsing failed` when the generated text is not valid JSON or does not match the expected structure.

Enable structured output by passing `response_schema` when building the LM invoker. `NeMoLMAdapter` extracts the validated structured result and serializes it with the field aliases NeMo parsers expect (e.g. `"User Safety"`, `"Safety Categories"`).

Define a Pydantic schema that matches the task output format in `config.yml`:

```python
from typing import Literal

from pydantic import BaseModel, ConfigDict, Field


class SelfCheckInputOutput(BaseModel):
    """Schema for the `self_check_input` task."""

    model_config = ConfigDict(populate_by_name=True, serialize_by_alias=True)

    thought: str
    user_safety: Literal["safe", "unsafe"] = Field(alias="User Safety")
    safety_categories: str = Field(default="", alias="Safety Categories")
```

Pass it to `build_lm_invoker`:

```python
from gllm_inference.builder import build_lm_invoker
from gllm_inference.schema.config import ThinkingConfig

invoker = build_lm_invoker(
    model_id="openai/gpt-5-nano",
    credentials=os.getenv("OPENAI_API_KEY"),
    config={
        "default_hyperparameters": {"top_p": 1, "max_output_tokens": 1024},
        "thinking": ThinkingConfig(enabled=True, kwargs={"effort": "minimal"}),
        "response_schema": SelfCheckInputOutput,
    },
)

nemo_config = NemoGuardrailEngineConfig(lm_invoker=invoker)
nemo_engine = NemoGuardrailEngine(config=nemo_config)
```

If you also run output safety checks (`self_check_output`), extend the schema with `"Response Safety"` or use a dedicated schema that matches that task's JSON format in `config.yml`.

**Notes:**

1. Set `serialize_by_alias=True` when field names in `config.yml` contain spaces (e.g. `"User Safety"`).
2. Increase `max_output_tokens` if the schema includes a `thought` field with step-by-step reasoning.
3. Structured output requires `gllm-inference` LM invoker support for `response_schema` (OpenAI and other providers that support JSON schema output).
