Metadata-Version: 2.4
Name: cjk-semantic-split
Version: 0.2.0
Summary: Recursive atomic decomposition of CJK characters with 9-grid 2-digit spatial encoding. 将汉字递归拆解为原子部件，并在 9 宫格 2 位数字空间编码中标记每个组件的位置。
Project-URL: Source, https://github.com/yueguanyu/Semantic-Character-Splitting-and-Structural-Coding-for-Chinese-Script.git
Project-URL: Bug Tracker, https://github.com/yueguanyu/Semantic-Character-Splitting-and-Structural-Coding-for-Chinese-Script/issues
Author-email: Guanyu Yue <yue.guanyu@hotmail.com>
License: MIT License
        
        Copyright (c) 2026
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: chinese,cjk,decomposition,ideograph
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.9
Provides-Extra: dev
Requires-Dist: hatch>=1.0; extra == 'dev'
Requires-Dist: hypothesis>=6.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.1; extra == 'dev'
Description-Content-Type: text/markdown

# cjk-semantic-split



---

## 中文简介

`cjk-semantic-split` 是一个 Python 库，用于将**任意 CJK 统一表意文字**递归拆分为原子部件，并使用 **9 宫格 2 位数字空间编码**标记每个组件的位置。

**核心特性**

- **递归拆解**：例如 `踩 → 足 + 采 → 爪 + 木`，中间非叶节点（如 `采`）完整保留
- **位置编码**：每个组件的 2 位数字代号（`00/10/20/30/40` 四个基本方位 + `50/60/70/80` 四个角方位 + `14/16/34/36` 四分拆壁 + `90–99` 包围框）
- **组合算子 `⊙`**：父子位置按空间几何合成（详见 [`docs/component-position-encoding.md`](docs/component-position-encoding.md)）
- **多重数据源**：内置四级 — Unihan kIDS / wikimedia Commons / 下一层采集 / 手工修补
- **零运行时依赖**：纯 Python

**主要 API**

- `decompose(text)` → 字符串（默认 JSON 输出）
- `read_structure(text)` → `list[DecompositionTree]`，保留完整嵌套层级（包含中间非叶节点）
- `python -m chinese_decompose <字符>` → 命令行

```python
from chinese_decompose import read_structure

trees = read_structure("踩")
# trees[0].atoms     == ('足', '爪', '木')
# trees[0].positions == (30, 60, 80)
# 采 作为中间节点保留在 children[1] 中
```

详细算法参见 [`docs/decomposition-logic.md`](docs/decomposition-logic.md)；完整位置编码表参见 [`docs/component-position-encoding.md`](docs/component-position-encoding.md)。

---

Recursive atomic decomposition of CJK characters with 9-grid spatial position encoding.

```bash
pip install cjk-semantic-split
```

```python
from chinese_decompose import decompose

decompose("踩")
# '踩 [30]:足 [60]:爪 [80]:木'

decompose("囚")
# '囚 [90]:囗 [0]:人'

decompose("踩好", as_tree=True)
# [DecompositionTree('踩'), DecompositionTree('好')]

# Read the full structural decomposition (preserves intermediate nodes
# like 采 inside 踩; each tree node carries its 2-digit position code).
from chinese_decompose import read_structure
trees = read_structure("踩")
trees[0].atoms     # ('足', '爪', '木')
trees[0].positions # (30, 60, 80)
trees[0].children[1].char  # '采'  (intermediate node preserved!)
trees[0].children[1].position  # 40
```

## Documentation

- [Component Position Encoding](docs/component-position-encoding.md) — full position-code reference (00/10/20/30/40, corners, 4-cell strips, surrounds) and the composition operator.
- [Decomposition Logic](docs/decomposition-logic.md) — algorithm walkthrough, data flow, worked example `踩 → 足 + 采 → 爪 + 木`.

## Position Encoding

Every component sits at a 2-digit position code: cardinals (`10/20/30/40`), single-cell corners (`50/60/70/80`), 4-cell layout strips (`14/16/34/36`), or surround envelopes (`90`–`99`). See the [design spec](docs/superpowers/specs/2026-07-16-cjk-semantic-split-design.md) for full details.

## CLI

```bash
python -m chinese_decompose 踩                  # JSON output (default)
python -m chinese_decompose 踩 --format dsl    # DSL string
python -m chinese_decompose 踩 --format tree   # repr() of nested trees
python -m chinese_decompose --self-test         # coverage report
```

## Data Sources

Decomposition data is bundled for offline use. Coverage extends across four tiers:

- Tier 1 (Unihan kIDS): ~21k chars
- Tier 2 (wikimedia Commons): ~22k chars
- Tier 3 (next-layer harvest): ~80k chars
- Tier 4 (manual patches): dispute-case overrides

Last-wins resolution: patches > l3 > l2 > l1.

## Configuration

```python
from chinese_decompose import set_primitives

set_primitives({"木", "火", "水", ...})  # process-local override
```

## License

MIT — see `LICENSE`.
