Metadata-Version: 2.5
Name: pyvizion
Version: 1.0.0
Summary: Automação de interface por visão computacional (imagem + OCR) para sistemas legados.
Author: Josias Azevedo da Silva
License: Copyright (c) 2026 Josias Azevedo da Silva. All rights reserved.
        
        This software is proprietary and confidential. Unauthorized copying,
        distribution, modification, or use of this software, via any medium,
        is strictly prohibited without prior written permission.
License-File: LICENSE
Keywords: automation,computer-vision,gui-automation,legacy-systems,ocr,opencv,oracle-forms,rpa,template-matching,tesseract
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: License :: Other/Proprietary License
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.9
Requires-Dist: mss>=9.0
Requires-Dist: numpy>=1.21
Requires-Dist: opencv-python-headless>=4.5
Requires-Dist: pillow>=9.0
Requires-Dist: pyautogui>=0.9.53
Requires-Dist: pyperclip>=1.8
Requires-Dist: pytesseract>=0.3.8
Requires-Dist: pyyaml>=6.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Description-Content-Type: text/markdown

# pyvizion

![Python](https://img.shields.io/badge/python-3.9+-blue.svg)
![Versão](https://img.shields.io/badge/vers%C3%A3o-1.0.0-green.svg)

**pyvizion** automatiza qualquer programa pela tela, como uma pessoa faria: encontra **imagens** (OpenCV, várias escalas) e **textos** (OCR com Tesseract), **clica**, **digita** e **espera** — com **backtrack** automático. Feito para **não falhar em sistemas legados**: ERPs, Oracle Forms, Delphi/VB6, Java Swing, Citrix/RDP e emuladores de terminal.

```python
from pyvizion import Vizion

vz = Vizion()
vz.click_image("button1.png", backtrack=True)
vz.click_text("Save", backtrack=True)       # se falhar, refaz o button1 e tenta de novo
vz.click_text("Confirm", backtrack=True)    # se falhar, refaz o Save e tenta de novo
```

---

## 📦 Instalação

```bash
pip install pyvizion
```

Uma instalação só, com tudo incluído e já no modo mais rápido.

**Tesseract** (só para funções de texto — `click_text`, `find_text`, `read_text`):

- **Windows:** <https://github.com/UB-Mannheim/tesseract/wiki> — instale em `C:\Program Files\Tesseract-OCR\` e marque **Portuguese**.
- **Linux:** `sudo apt-get install tesseract-ocr tesseract-ocr-por`
- **macOS:** `brew install tesseract tesseract-lang`

```bash
python -m pyvizion doctor      # confere dependências, Tesseract, idiomas e monitores
```

---

## 🎯 Uso rápido

### Backtrack entre métodos

```python
vz = Vizion()
vz.click_image("menu.png", backtrack=True)
vz.click_text("Relatórios", backtrack=True)   # falhou? reexecuta o anterior e tenta de novo
```

### Sessão de backtrack

```python
vz.start_task_session()
vz.click_image("button1.png", backtrack=True)
vz.click_text("Clientes", backtrack=True)
vz.click_text("Novo", backtrack=True)
successful, total = vz.end_task_session()
print(f"Sucesso: {successful}/{total}")
```

### Lista de tarefas

```python
from pyvizion import execute_tasks

tasks = [
    {"image": "button.png", "region": (100, 100, 200, 50), "confidence": 0.9,
     "specific": False, "backtrack": True, "delay": 1, "mouse_button": "left"},

    {"text": "Login", "region": (50, 50, 300, 100), "char_type": "letters",
     "backtrack": True, "sendtext": "usuario123{tab}senha{enter}"},

    {"type": "relative_image", "anchor_image": "warning_icon.png",
     "target_image": "ok_button.png", "max_distance": 200},

    {"type": "click", "x": 500, "y": 300, "mouse_button": "right"},
    {"type": "type_text", "text": "Hello World!"},
    {"type": "keyboard_command", "command": "Ctrl+S", "delay": 1},
]

execute_tasks(tasks)            # ou: Vizion().execute_tasks(tasks)
```

---

## 🔧 Métodos

| Método | O que faz |
|---|---|
| `click_image(image_path, region, confidence, delay, mouse_button, max_attempts, backtrack, specific, sendtext, show_overlay)` | Encontra e clica numa imagem |
| `find_image(image_path, region, confidence, max_attempts, backtrack, specific, scales)` | Retorna `(x, y, largura, altura)` ou `None` |
| `click_text(text, region, filter_type, delay, mouse_button, occurrence, backtrack, max_attempts, sendtext, confidence_threshold, show_overlay)` | Encontra e clica num texto |
| `find_text(text, region, filter_type, confidence_threshold, occurrence, max_attempts, backtrack)` | Caixa do texto ou `None` |
| `click_relative_image(anchor_image, target_image, max_distance, confidence, target_region, delay, mouse_button, backtrack, max_attempts)` | Clica no alvo mais próximo de uma âncora |
| `find_relative_image(...)` | Caixa do alvo mais próximo da âncora |
| `click_image_near_text(anchor_text, target_image, ...)` | Clica na imagem mais próxima de um texto |
| `click_coordinates(x, y, delay, mouse_button, backtrack)` / `click_at(location, ...)` | Clique em coordenadas |
| `type_text(text, interval, delay, backtrack)` | Digita texto (aceita macros) |
| `keyboard_command(command, delay, backtrack)` | `"Ctrl+S"`, `"F7"`, `"Alt+Tab"`... |
| `wait_for_image(...)` / `wait_for_text(...)` / `wait_until_gone(...)` | Esperas |
| `image_exists(...)` / `text_exists(...)` / `find_all_images(...)` | Verificações |
| `click_any([...])` / `find_any([...])` | Primeiro alvo encontrado dentre alternativas |
| `read_text(region, single_line)` | Lê o texto de uma área |
| `focus_window(title)` / `wait_window(title)` / `window_region(title)` | Janelas |
| `press(key, presses)` / `hotkey(*keys)` / `scroll(clicks)` / `drag(start, end)` / `screenshot(path)` | Utilidades |
| `execute_tasks(tasks)` / `execute_with_backtrack_between_tasks(tasks)` | Listas de tarefas |
| `start_task_session()` / `end_task_session()` | Sessão de backtrack |
| `configure_overlay(enabled, color, duration, width)` / `get_overlay_config()` / `test_overlay_colors()` | Overlay visual |
| `list_windows()` | Títulos das janelas abertas |

**Parâmetros comuns**

- `region=(x, y, largura, altura)` — onde procurar. Use sempre que puder (OCR ~6× mais rápido, sem falsos positivos). `vz.window_region("Título")` devolve a área de uma janela.
- `specific=True` — busca só na `region`. `specific=False` — tenta a região e depois a tela inteira, em escalas 0,8× a 1,2×.
- `mouse_button` — `"left"`, `"right"`, `"double"`, `"move_to"` (só passa o mouse), `"middle"`, `"triple"`.
- `filter_type` — `"letters"`, `"numbers"` ou `"both"`. `occurrence=2` — a segunda ocorrência em ordem de leitura.
- `offset=(dx, dy)` — clica deslocado do alvo (ex.: o campo à direita do rótulo).
- `timeout=` — espera o alvo aparecer por até N segundos.

Os métodos também existem como funções, sem instância: `from pyvizion import click_image, find_text, execute_tasks`.

---

## 📝 Preenchendo campos (`sendtext`)

`sendtext` é o texto digitado **logo depois do clique**. Dentro dele, palavras entre chaves `{ }` são **teclas**:

```python
vz.click_text("Usuário", offset=(150, 0), sendtext="admin{tab}senha123{enter}")
```

O que acontece, em ordem:

| # | Trecho | O pyvizion faz |
|---|---|---|
| 1 | *(clique)* | clica 150 px à direita do rótulo "Usuário" (dentro do campo) |
| 2 | `admin` | digita "admin" |
| 3 | `{tab}` | aperta **Tab** → vai para o campo de senha |
| 4 | `senha123` | digita "senha123" |
| 5 | `{enter}` | aperta **Enter** → confirma |

### Receitas do dia a dia

| Quero... | `sendtext` |
|---|---|
| digitar num campo vazio | `"12345"` |
| **substituir** o que já está no campo | `"{ctrl}a{del}12345"` |
| digitar e ir para o próximo campo | `"12345{tab}"` |
| digitar e confirmar | `"12345{enter}"` |
| preencher dois campos seguidos | `"01/01/2026{tab}31/01/2026"` |
| pular campos | `"{tab*3}"` |
| esperar o sistema reagir | `"12345{tab}{wait 1}"` |

Prefira **um passo por campo** — fica mais fácil de ler e, se algo falhar, o log mostra exatamente onde:

```python
vz.click_text("Data Inicial", offset=(150, 0), sendtext="{ctrl}a{del}01/01/2026")
vz.click_text("Data Final",   offset=(150, 0), sendtext="{ctrl}a{del}31/01/2026")
vz.keyboard_command("Enter")
```

### Teclas disponíveis

| Macro | Tecla |
|---|---|
| `{enter}` `{tab}` `{esc}` `{space}` | Enter, Tab, Esc, Espaço |
| `{del}` `{backspace}` `{insert}` | Delete, Backspace, Insert |
| `{up}` `{down}` `{left}` `{right}` | setas |
| `{home}` `{end}` `{pgup}` `{pgdn}` | Home, End, Page Up, Page Down |
| `{f1}` … `{f12}` | teclas de função |
| `{ctrl}a` `{alt}f` `{shift}x` | modificador + a próxima letra |
| `{ctrl+shift+s}` `{alt+f4}` | combinação completa |
| `{tab*3}` `{down 5}` | repetição |
| `{wait 1.5}` `{wait 300ms}` | pausa |
| `{{` `}}` | as chaves `{` e `}` literais |

- Maiúsculas não importam: `{Enter}` = `{enter}`.
- Chaves com algo que não é tecla são digitadas normalmente: `"valor {total}"` digita exatamente isso.
- Acentos e `ç` saem corretos, e a área de transferência do usuário é restaurada.
- Em **Citrix/RDP/terminais** que não aceitam colar, use `Vizion({"typing_mode": "type"})`.

Para apertar **só uma tecla**, sem texto: `vz.keyboard_command("Ctrl+S")`, `vz.press("tab", 3)`, `vz.hotkey("ctrl", "shift", "s")`.

---

## 📋 Tarefas: todos os tipos

```python
tasks = [
    {"type": "focus_window", "title": "Oracle Applications"},
    {"type": "wait_image", "image": "tela_principal.png", "timeout": 30},
    {"text": "Cliente", "offset": (160, 0), "sendtext": "12345{enter}", "required": True},
    {"type": "wait_image", "image": "ampulheta.png", "gone": True, "timeout": 60},
    {"type": "relative_image", "anchor_text": "Pedido", "target_image": "lupa.png"},
    {"image": "popup_aviso.png", "optional": True, "sendtext": "{enter}"},
    {"type": "wait_text", "text": "Registro salvo", "timeout": 15},
    {"type": "wait", "seconds": 1},
    {"type": "scroll", "clicks": -5},
]
```

| Tipo | Chaves |
|---|---|
| imagem | `image`, `region`, `confidence`, `specific`, `scales` |
| texto | `text`, `region`, `char_type`, `occurrence`, `confidence_threshold` |
| `relative_image` | `anchor_image` ou `anchor_text`, `target_image`, `max_distance`, `target_region` |
| `click` | `x`, `y` |
| `type_text` | `text`, `interval` |
| `keyboard_command` | `command` |
| `wait_image` / `wait_text` | `image`/`text`, `timeout`, `gone` |
| `wait` · `focus_window` · `scroll` | `seconds` · `title` · `clicks`, `x`, `y` |

Chaves de qualquer tarefa: `mouse_button`, `delay`, `sendtext`, `offset`, `backtrack`, `max_attempts`, `timeout`, `optional` (falha não conta nem faz backtrack), `required` (falha interrompe a lista), `show_overlay`, `click_hold`.

---

## ⚙️ Configuração

```python
vz = Vizion({
    "confidence_threshold": 80.0,
    "tesseract_lang": "por",
    "show_overlay": False,
    "image_folders": ["./imagens"],
    "save_failure_screenshots": True,
})

vz.config.set("show_overlay", True)          # alterar depois
vz = Vizion("pyvizion.json")                  # ou de um arquivo .json/.yaml
```

| Chave | Padrão | Descrição |
|---|---|---|
| `confidence_threshold` | `75.0` | Limiar do OCR (0–100) |
| `default_confidence` | `0.9` | Confiança padrão de imagens nas tarefas |
| `tesseract_path` / `tessdata_path` | auto | Caminhos do Tesseract |
| `tesseract_lang` | auto (`por` se instalado) | Idioma(s): `"por"`, `"por+eng"` |
| `image_processing_methods` | `"all"` | `"all"`, `"balanced"`, `"fast"` ou lista de técnicas |
| `ocr_large_image_methods` | `"fast"` | Técnicas usadas na tela inteira |
| `preprocessing_enabled` | `True` | `False` = OCR só na imagem original |
| `ocr_upscale` / `ocr_workers` | `2.0` / auto | Ampliação de áreas pequenas / paralelismo |
| `ocr_fuzzy` / `ocr_fuzzy_threshold` | `True` / `0.8` | Tolerância a erros do OCR |
| `grayscale` · `scales` · `dpi_scales` · `edge_fallback` | `True` · 0,8–1,2 · `True` · `False` | Busca de imagem |
| `min_confidence` · `retry_delay` | `0.7` · `0.5` | Tentativas |
| `overlay_enabled` | `True` | Liga/desliga o sistema de overlay |
| `image_folders` | `[]` | Pastas onde procurar imagens |
| `typing_mode` · `restore_clipboard` · `typing_interval` | `"paste"` · `True` · `0.02` | Digitação |
| `click_hold` · `move_pause` · `movement_duration` | `0` · `0.15` · `0.1` | Ritmo do mouse |
| `failsafe` | `True` | Mouse no canto superior esquerdo interrompe |
| `show_overlay` · `overlay_color` · `overlay_duration` · `overlay_width` | `False` · `red` · `1000` · `4` | Retângulo sobre o alvo antes do clique (depuração) |
| `save_failure_screenshots` · `failure_screenshot_dir` | `False` · `pyvizion_failures` | Print a cada falha |
| `stop_on_failure` · `max_backtrack_attempts` | `False` · `2` | Listas de tarefas |
| `log_level` | `"INFO"` | |

---

## 🛡️ Robustez para sistemas legados

- **DPI e vários monitores** — modo DPI por monitor, captura de todos os monitores (inclusive coordenadas negativas); a escala de cada monitor (125%, 150%) entra automaticamente na busca.
- **OCR tolerante** — ignora acentos e maiúsculas, corrige `0/O`, `1/l`, `5/S`, aceita pequenas diferenças e palavras coladas/quebradas, amplia áreas pequenas, inverte texto claro em fundo escuro, roda em paralelo e usa consenso entre leituras.
- **Captura rápida** — ~5 ms por área.
- **Esperas**, **alternativas** (`click_any`), **âncoras** (`click_relative_image`, `click_image_near_text`), **clique deslocado** (`offset`), **leitura de campos** (`read_text`), **foco de janela**.
- **Listas seguras** — `required`, `optional`, `stop_on_failure`.
- **Diagnóstico** — prints de falha e `python -m pyvizion doctor`.

---

## 🛠️ Linha de comando

```bash
python -m pyvizion doctor               # ambiente
python -m pyvizion screenshot tela.png  # print para recortar suas imagens
python -m pyvizion position             # x/y do mouse em tempo real (Ctrl+C)
python -m pyvizion pick                 # marca região (2 cantos), OCR e trechos para colar no código
python -m pyvizion ocr 100 200 300 40   # lê o texto de uma área (se já souber x,y,w,h)
```

---

## 💡 Dicas

1. Use **`region`** (ou `vz.window_region("Título")`) sempre que puder.
2. Recorte **imagens pequenas e únicas** (o ícone, não o botão inteiro), em PNG.
3. Prefira **`wait_until_gone("ampulheta.png")`** a `delay` fixo.
4. Comece com **`focus_window("Título")`**.
5. Instale o idioma **`por`** do Tesseract.
6. Para parar o robô, leve o mouse ao **canto superior esquerdo** da tela.

---

## 🔄 Migrando de `bot-vision-suite` / `visus-desktop`

Substituição mínima:

```diff
- pip install bot-vision-suite
+ pip install pyvizion

- from bot_vision import BotVision
+ from pyvizion import Vizion

- bot = BotVision()
+ vz = Vizion()
```

Funções sem instância continuam iguais em espírito (`from pyvizion import click_image, execute_tasks, …`).

### Tabela 1:1 (métodos)

| bot-vision-suite / `BotVision` | pyvizion / `Vizion` | Observação |
|---|---|---|
| `BotVision()` | `Vizion()` ou `Vizion({...})` / `Vizion("config.yaml")` | pyvizion aceita dict, JSON ou YAML |
| `BotVision(config=dict)` | `Vizion(config=dict)` + `Config` | Chaves parecidas; veja [Configuração](#-configuração) |
| `click_image(...)` | `click_image(...)` | **Compatível** |
| `find_image(...)` | `find_image(...)` | Retorno: `(x, y, w, h)` ou `None` (mesma ideia) |
| `find_all_images(...)` | `find_all_images(...)` | **Compatível** |
| `image_exists(...)` | `image_exists(...)` | **Compatível** |
| `click_text(...)` | `click_text(...)` | **Compatível** |
| `find_text(...)` | `find_text(...)` | **Compatível** |
| `text_exists(...)` | `text_exists(...)` | **Compatível** |
| `read_text(region, single_line)` | `read_text(region, single_line)` | **Compatível** |
| `click_relative_image(anchor, target, ...)` | `click_relative_image(...)` | **Compatível** |
| `find_relative_image(...)` | `find_relative_image(...)` | **Compatível** |
| `click_image_near_text(anchor_text, target, ...)` | `click_image_near_text(...)` | **Compatível** |
| `click_coordinates(x, y, ...)` | `click_coordinates(x, y, ...)` | **Compatível** |
| `click_at(location, ...)` | `click_at(location, ...)` | **Compatível** |
| `type_text(...)` | `type_text(...)` | Macros `{enter}`, `{tab}`, `{ctrl}a`… |
| `keyboard_command(...)` | `keyboard_command(...)` | **Compatível** |
| `wait_for_image(...)` | `wait_for_image(...)` | **Compatível** |
| `wait_for_text(...)` | `wait_for_text(...)` | **Compatível** |
| `wait_until_gone(...)` | `wait_until_gone(image_path=..., text=...)` | **Compatível** |
| `click_any([...])` | `click_any([...])` | pyvizion retorna **índice** clicado ou `None` (não só bool) |
| `find_any([...])` | `find_any([...])` | Retorno `(índice, region)` ou `(None, None)` |
| `execute_tasks(tasks)` | `execute_tasks(tasks)` | **Compatível**; pyvizion tem ainda `optional`, `required`, `stop_on_failure` |
| `execute_with_backtrack_between_tasks(...)` | `execute_with_backtrack_between_tasks(...)` | Formato `{'type', 'params'}` |
| `start_task_session()` / `end_task_session()` | idem | Retorno `(ok, total)` |
| `configure_overlay(...)` | `configure_overlay(...)` | **Compatível** |
| `get_overlay_config()` | `get_overlay_config()` | **Compatível** |
| `test_overlay_colors()` | `test_overlay_colors()` | **Compatível** |
| `focus_window` / `wait_window` / `window_region` | idem | **Compatível** |
| `list_windows()` | `list_windows()` | **Compatível** |
| `press` / `hotkey` / `scroll` / `drag` / `screenshot` | idem em `Vizion` | **Compatível** |
| `from bot_vision import click_image, …` | `from pyvizion import click_image, …` | Mesmo padrão |
| `python -m bot_vision doctor` (se existir) | `python -m pyvizion doctor` | + `screenshot`, `position`, `ocr` |

### Parâmetros que mudam de nome (tarefas / dicts)

| BVS / visus | pyvizion | Notas |
|---|---|---|
| `"filter_type": "letters"` em tarefas | `"char_type": "letters"` | Nos **métodos** continua `filter_type=` |
| `"specific": false` (busca tela inteira + escalas) | `"specific": false` | Internamente vira `flexible=True` na busca |
| `"anchor_image"` + `"target_image"` | idem | + `"anchor_text"` em `type: relative_image` |
| `"confidence_threshold"` (OCR) | `"confidence_threshold"` ou `"early_confidence"` | Nos métodos: `confidence_threshold=` |

### O que o pyvizion **adiciona** (não existe na BVS da mesma forma)

| Recurso | Onde |
|---|---|
| Config em **JSON/YAML** + `config.set()` com reload | `Vizion("projeto.json")` |
| Pastas de imagens (`image_folders`) | Config |
| Screenshot automático em falha | `save_failure_screenshots` |
| OCR multi-técnica nomeada (`fast` / `balanced` / `all`) | `image_processing_methods` |
| DPI por monitor + captura multi-monitor | import + `dpi_scales` |
| Matching por **bordas** (tema claro/escuro) | `edge_fallback` |
| Modo digitação para Citrix/RDP | `typing_mode: "type"` |
| Alvos tipados (`Image`, `Text`, `Near`, `AnyOf`) | `core/specs.py` (API interna/moderna) |
| Tarefas `optional` / `required` / `click_hold` | `execute_tasks` |

### O que **ainda falta** para drop-in 100% com `BotVision`

| Item | Situação | Workaround hoje |
|---|---|---|
| **`BotVision` alias** | Não exportado | `from pyvizion import Vizion as BotVision` |
| **`register_image("ok", "ok.png")`** (visus-desktop) | Não implementado | Caminho completo ou `image_folders: ["./imagens"]` |
| **`limpar_texto` / `clean_text`** | Não implementado | Normalizar string antes do OCR manualmente |
| **Pacote `bot_vision`** | Nome diferente | Só trocar imports para `pyvizion` |
| **EasyOCR / PyTorch** (opcional na BVS) | Não incluso | Só Tesseract (proposital — mais leve) |
| **Extras AI (`openai`)** da BVS | Não incluso | Fora do escopo desktop |
| **Automação web (Selenium/DOM)** | Não incluso | BVS também não tem; use Selenium à parte |

Com a troca de import e de `BotVision` → `Vizion`, a maior parte dos scripts BVS roda **sem alterar** chamadas de `click_image`, `click_text`, `execute_tasks` e backtrack.

---

## 🧪 Desenvolvimento

```bash
pip install -e .[dev]
pytest
```

## 📄 Licença

Software proprietário — veja [LICENSE](LICENSE).

Desenvolvido por **Josias Azevedo da Silva**.
