Metadata-Version: 2.4
Name: scalellm-ai
Version: 0.3.0
Summary: A simple high-level AI toolkit with scratch-trained LLM, image, video, and TTS models plus easy pretrained wrappers.
License: MIT License
        
        Copyright (c) 2026 ScaleLLM AI contributors
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
Classifier: Programming Language :: Python :: 3
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch>=2.2
Requires-Dist: tiktoken>=0.7
Requires-Dist: numpy>=1.24
Provides-Extra: text
Requires-Dist: transformers>=4.45; extra == "text"
Requires-Dist: accelerate>=1.0; extra == "text"
Provides-Extra: image
Requires-Dist: diffusers>=0.35; extra == "image"
Requires-Dist: transformers>=4.45; extra == "image"
Requires-Dist: accelerate>=1.0; extra == "image"
Requires-Dist: safetensors>=0.4; extra == "image"
Requires-Dist: Pillow>=10; extra == "image"
Provides-Extra: video
Requires-Dist: diffusers>=0.35; extra == "video"
Requires-Dist: transformers>=4.45; extra == "video"
Requires-Dist: accelerate>=1.0; extra == "video"
Requires-Dist: safetensors>=0.4; extra == "video"
Requires-Dist: imageio>=2.35; extra == "video"
Requires-Dist: imageio-ffmpeg>=0.5; extra == "video"
Provides-Extra: tts
Requires-Dist: transformers>=4.45; extra == "tts"
Requires-Dist: accelerate>=1.0; extra == "tts"
Requires-Dist: scipy>=1.11; extra == "tts"
Requires-Dist: phonemizer>=3.2; extra == "tts"
Provides-Extra: game
Requires-Dist: stable-baselines3>=2.5; extra == "game"
Requires-Dist: gymnasium>=1.0; extra == "game"
Provides-Extra: vision
Requires-Dist: transformers>=4.45; extra == "vision"
Requires-Dist: accelerate>=1.0; extra == "vision"
Requires-Dist: Pillow>=10; extra == "vision"
Provides-Extra: speech
Requires-Dist: transformers>=4.45; extra == "speech"
Requires-Dist: accelerate>=1.0; extra == "speech"
Requires-Dist: librosa>=0.10; extra == "speech"
Requires-Dist: soundfile>=0.12; extra == "speech"
Provides-Extra: audio
Requires-Dist: transformers>=4.45; extra == "audio"
Requires-Dist: accelerate>=1.0; extra == "audio"
Requires-Dist: scipy>=1.11; extra == "audio"
Provides-Extra: embeddings
Requires-Dist: sentence-transformers>=3.2; extra == "embeddings"
Provides-Extra: all
Requires-Dist: transformers>=4.45; extra == "all"
Requires-Dist: accelerate>=1.0; extra == "all"
Requires-Dist: diffusers>=0.35; extra == "all"
Requires-Dist: safetensors>=0.4; extra == "all"
Requires-Dist: Pillow>=10; extra == "all"
Requires-Dist: imageio>=2.35; extra == "all"
Requires-Dist: imageio-ffmpeg>=0.5; extra == "all"
Requires-Dist: scipy>=1.11; extra == "all"
Requires-Dist: phonemizer>=3.2; extra == "all"
Requires-Dist: stable-baselines3>=2.5; extra == "all"
Requires-Dist: gymnasium>=1.0; extra == "all"
Requires-Dist: librosa>=0.10; extra == "all"
Requires-Dist: soundfile>=0.12; extra == "all"
Requires-Dist: sentence-transformers>=3.2; extra == "all"
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: twine>=5; extra == "dev"
Requires-Dist: pytest>=8; extra == "dev"
Dynamic: license-file

# ScaleLLM AI

`scalellm-ai` is a tiny high-level Python toolkit for building AI applications with very little code.

It now supports **scratch-trained models** for:
- LLMs
- image generation
- video generation
- text-to-speech

It also still includes easy wrappers around strong **pretrained** models for images, video, TTS, vision, speech recognition, embeddings, text generation, and zero-shot classification.

## Important reality check

The scratch-trained image, video, and TTS models are **real trainable models from random weights**, but they are **toy-sized starter architectures** meant for learning and experimentation.

That means:
- you can genuinely train them from scratch
- they can learn small custom datasets
- they will not match Stable Diffusion, Sora, or commercial TTS quality
- they are designed so the **user-facing code stays under 20 lines**

## Install

Core package:

```bash
pip install scalellm-ai
```

Feature installs:

```bash
pip install "scalellm-ai[image]"
pip install "scalellm-ai[video]"
pip install "scalellm-ai[tts]"
pip install "scalellm-ai[game]"
pip install "scalellm-ai[all]"
```

## 1. Train your own tiny LLM from scratch

```python
from scalellm_ai import ScaleLLM

llm = ScaleLLM(d_model=128, n_layer=4, n_head=4, max_seq_len=128)
text = open("dataset.txt", encoding="utf-8").read()
llm.train(text, steps=500, batch_size=4)
print(llm.generate("Python is", max_new_tokens=80))
llm.save("my_llm.pt")
```

## 2. Train an image model from scratch

Manifest format:

```text
images/cat1.png|orange cat
images/car1.png|red sports car
```

```python
from scalellm_ai import ImageAI

ai = ImageAI.from_scratch(image_size=32)
ai.train("image_manifest.txt", epochs=20)
ai.generate("red sports car", output="car.png")
ai.save("my_image_ai.pt")
```

## 3. Train a video model from scratch

Manifest format:

```text
clips/fly.gif|bird flying
clips/wave.gif|ocean waves
```

```python
from scalellm_ai import VideoAI

video = VideoAI.from_scratch(image_size=32, num_frames=8)
video.train("video_manifest.txt", epochs=20)
video.generate("bird flying", output="bird.gif", fps=6)
video.save("my_video_ai.pt")
```

## 4. Train a TTS model from scratch

Manifest format:

```text
hello.wav|hello world
welcome.wav|welcome to scale llm ai
```

```python
from scalellm_ai import TTS

voice = TTS.from_scratch(sample_rate=8000, seconds=1.0)
voice.train("tts_manifest.txt", epochs=20)
voice.speak("hello world", output="hello.wav")
voice.save("my_tts.pt")
```

## 5. Pretrained image generation

```python
from scalellm_ai import ImageAI

ai = ImageAI()
ai.generate("cinematic robot exploring an ancient library", output="robot.png")
```

## 6. Pretrained video generation

```python
from scalellm_ai import VideoAI

video = VideoAI()
video.generate("a small spacecraft flying through glowing blue clouds", output="space.mp4")
```

## 7. Pretrained TTS

```python
from scalellm_ai import TTS

voice = TTS()
voice.speak("ScaleLLM AI can turn this sentence into speech.", output="speech.wav")
```

## 8. Game-playing AI

```python
from scalellm_ai import GameAI

agent = GameAI("CartPole-v1", algorithm="PPO", verbose=1)
agent.train(20_000)
agent.save("cartpole_agent")
```

## 9. Vision AI

```python
from scalellm_ai import VisionAI

vision = VisionAI()
for result in vision.classify("photo.jpg", top_k=3):
    print(result["label"], result["score"])
```

## 10. Speech-to-text

```python
from scalellm_ai import SpeechAI

speech = SpeechAI()
print(speech.transcribe("meeting.wav"))
```

## 11. Embeddings / semantic similarity

```python
from scalellm_ai import EmbedAI

embed = EmbedAI()
print(embed.similarity("A dog is running outside.", "A puppy is playing outdoors."))
```

## 12. Pretrained text generation

```python
from scalellm_ai import TextAI

ai = TextAI()
print(ai.generate("Explain neural networks simply:", max_new_tokens=120))
```

## 13. Zero-shot classifier

```python
from scalellm_ai import ClassifyAI

ai = ClassifyAI()
result = ai.classify("The team won the championship last night.", ["sports", "business", "science"])
print(result["labels"][0])
```

## 14. Music / text-to-audio

```python
from scalellm_ai import MusicAI

music = MusicAI()
music.generate("energetic retro arcade synthwave with a heroic melody", output="theme.wav")
```

## 15. Object detection

```python
from scalellm_ai import ObjectAI

ai = ObjectAI()
for obj in ai.detect("street.jpg", threshold=0.7):
    print(obj["label"], obj["score"], obj["box"])
```

## Notes on scratch datasets

- **ImageAI.from_scratch** expects a manifest with `image_path|caption`
- **VideoAI.from_scratch** expects a manifest with `video_path|caption`
- **TTS.from_scratch** expects a manifest with `audio_path|text`
- all scratch models start from **random weights**
- all scratch models are kept intentionally small so they can be understood and extended

## Examples

Every file in `examples/` is intentionally 20 lines or fewer.

## Publishing safely

Install the publishing tools first with `pip install "scalellm-ai[dev]"`.

Never put PyPI tokens in `publish.py`. Set the token in your shell instead:

PowerShell:

```powershell
$env:PYPI_API_TOKEN="pypi-..."
python publish.py
```

macOS/Linux:

```bash
export PYPI_API_TOKEN="pypi-..."
python publish.py
```

## License

MIT. Individual pretrained models can have their own licenses and usage terms; check model cards before redistribution or commercial use.
