Metadata-Version: 2.4
Name: simready-search
Version: 2026.7.1
Summary: SimReady Asset Search Library
Author: NVIDIA Corporation
License-Expression: Apache-2.0
Project-URL: Homepage, https://www.nvidia.com
Keywords: nvidia,simready,search
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE.txt
Requires-Dist: requests>=2.33.0
Requires-Dist: boto3==1.42.58
Requires-Dist: urllib3<3,>=2.7.0
Dynamic: license-file

# Asset Search Library (`simready.search`)

Python library for loading a SimReady project's asset metadata cache and performing filter-based searches.

## Initialization

Construct an instance of the `AssetLibrary` class.
Use one or more of these async methods to add a search source:
- `add_cache_source` for local project_config.toml
- `add_s3_source` for S3 project_config.toml
- `add_service_source` for web-hosted search service
- `add_indexed_source` for either local or S3 directories without caches
  - Only the `SearchFilterPathContains` will work to find files indexed this way, since the indexing will discover only file paths, no metadata.

```python
import asyncio
from simready.search import (
    AssetLibrary,
    AssetLibraryNetworkError,
    SearchFilterArbitraryDictValue,
    SearchFilterClass,
    SearchFilterCountry,
    SearchFilterFeature,
    SearchFilterHeight,
    SearchFilterLLMQuery,
    SearchFilterPathContains,
    SearchFilterScenePOI,
)

async def main():
    # Option 1: initialize with project_config.toml referring to a workspace_cache.json
    project_config_path = "d:/simready_foundations/sample_content/project_config.toml"
    asset_library = AssetLibrary()
    await asset_library.add_cache_source(project_config_path)

    # Option 2: no project_config, just index some directories
    asset_library = AssetLibrary()
    await asset_library.add_indexed_source("d:/some_content/", "local")
    await asset_library.add_indexed_source("d:/some_more_content/", "local")

    # Option 3: strict network error handling
    asset_library = AssetLibrary(raise_on_network_error=True)

asyncio.run(main())
```

`AssetLibrary` constructor args:

- `log_func` (default `None`): callback for library log messages.
- `raise_on_network_error` (default `False`): when `True`, raises `AssetLibraryNetworkError` for service/S3 failures.

When `raise_on_network_error=False`, network failures are logged and the library returns safe empty/partial results.

## Searching

`AssetLibrary.search(...)` takes 4 keyword arguments, each an optional list of search filters:

- **include_all**: include an asset if it passes **all** of these filters
- **include_any**: include an asset if it passes **any one** of these filters
- **exclude_all**: exclude an asset if it passes **all** of these filters
- **exclude_any**: exclude an asset if it passes **any one** of these filters

It returns a `SearchResults` sequence of `AssetData` for assets which were included and
not excluded. Iterate, index, and take `len` as you would a list. Per-search facts live on
`matches.search_metadata` -- see [Per-search metadata](#per-search-metadata).

## Result metadata

Metadata comes back at two levels, and the two never mix.

### Per-asset metadata

`AssetData.metadata` is an `AssetMetadata` -- the same type in every search mode -- or `None` when
the source reports nothing about that particular asset. It carries `relevance_score`, a float in
`[0, 1]` derived from the hit's raw search scores. Only service sources populate it today; cache
and indexed results report `None`.

### Per-search metadata

`matches.search_metadata` is a `SearchMetadata` describing the search itself: values
shared by every result, such as how a natural-language query was parsed. It is one dataclass for
every search mode, and every field is optional, because which ones are set depends on what kind of
source answered. Only service sources fill anything in today: the `raw_search_query` that was sent,
and the `llm_parse_result` described under
[Natural-language queries](#natural-language-queries).

A cache-only or index-only search -- and a search no source answered at all -- still sets
`search_metadata` to an empty `SearchMetadata`, so callers never have to check for `None` first.

## Examples

### Example 1: find really tall props which are neither structures nor trees

```python
matches = asset_library.search(
    include_all=[SearchFilterClass("prop"), SearchFilterHeight(minimum=25)],
    exclude_any=[SearchFilterPathContains("structures"), SearchFilterPathContains("vegetation")],
)
for asset_data in matches:
    print(asset_data.asset_path)
```

### Example 2: find red American signs

```python
matches = asset_library.search(
    include_all=[
        SearchFilterArbitraryDictValue(["factory", "sign", "panel_color"], "red"),
        SearchFilterCountry("usa"),
    ]
)
for asset_data in matches:
    print(asset_data.asset_path)
```

### Example 3: find people (or people-height objects), except for Chinese signs

```python
matches = asset_library.search(
    include_any=[SearchFilterClass("character"), SearchFilterHeight(minimum=1.55, maximum=1.95)],
    exclude_all=[SearchFilterCountry("china"), SearchFilterClass("sign")],
)
for asset_data in matches:
    print(asset_data.asset_path)
```

## Querying all values for a filter

### Example 4: query all possible values for Scene Point Of Interest (POI) tags

```python
poi_tags = asset_library.get_all_values(SearchFilterScenePOI)
print(sorted(poi_tags))
```

### Example 4a: find scenes that contain both parking and speed bumps

```python
matches = asset_library.search(
    include_all=[SearchFilterScenePOI("parking"), SearchFilterScenePOI("speedbump")],
)
for asset_data in matches:
    print(asset_data.asset_path)
```

## Feature validation filter

### Example 5: find assets validated to have feature `FET_005` with version `1.0.0`

```python
matches = asset_library.search(
    include_all=[SearchFilterFeature("FET_005", version="1.0.0")],  # version is optional
)
for asset_data in matches:
    print(asset_data.asset_path)
```

## Natural-language queries

`SearchFilterLLMQuery` sends a complex free-text query to the search service, which parses it into
a description plus structured filters. This lets one string express what would otherwise require
several filters, including constraints the library has no filter for.

Requires a service source whose deployment exposes the `/llm_parse/query` route.

### Example 6: describe the query instead of building it from filters

```python
matches = asset_library.search(
    include_all=[SearchFilterLLMQuery("robot with Isaac profile")],
)
search_metadata = matches.search_metadata
for asset_data in matches:
    print(f"{asset_data.metadata.relevance_score:.2f}: {asset_data.asset_path}")
```

The service interprets that as the description `"robot"` plus a filter on the validation profile.

### Seeing how the query was interpreted

A natural-language query can change the search in ways you did not ask for, so the parse is
reported back once for the whole search in `search_metadata.llm_parse_result`. The constraints the
service understood but could not enforce are the ones worth checking -- without them, a query
asking for a license or a file size looks like it worked while that clause was quietly ignored:

```python
parse = search_metadata.llm_parse_result or {}

for f in parse.get("interpreted_query", {}).get("filters", []):
    print(f"filtered on {f['field']} {f['operator']} {f['value']}")
for constraint in parse.get("unmapped_constraints", []):
    print(f"NOT enforced: {constraint['text']} -- {constraint['note']}")
```

For `"warehouse shelves with rigid body physics under 100MB"` against a deployment that indexes
neither physics nor file size, that prints two `NOT enforced` lines and no filters: the search was
really just `"warehouse shelves"`.

### How it combines with other filters

A parsed query **replaces** the other include filters for the service request, so express
constraints in the query text rather than alongside it:

- Other filters in `include_all` / `include_any` are ignored for the service request, and each one
  is named in a warning.
- `exclude_all` / `exclude_any` still apply. Exclusions are a separate axis the parsed query may
  not cover.
- `SearchFilterRelevance` still applies, because it filters the returned results rather than the
  request.
- `SearchFilterPhrase` is dropped in favor of the parsed query, with a warning, even when the two
  filters are in different include lists.
- Cache and indexed sources are unaffected and keep evaluating every filter normally. Since
  `SearchFilterLLMQuery` never matches a local asset, put it in `include_any` when you want local
  sources to contribute to the same search.
- `SearchFilterLLMQuery` and `SearchFilterPhrase` are only valid in `include_all` / `include_any`.
  Passing either in an exclude list is ignored with a warning: a parsed query or phrase cannot be
  negated.

If the deployment cannot parse queries, the query text degrades to a plain `SearchFilterPhrase`
and the other filters are applied as usual, so the search still returns useful results.

Constraints the service understood but could not enforce are reported as warnings through
`log_func` rather than failing the search.
