Metadata-Version: 2.5
Name: pyvizion
Version: 1.0.2
Summary: Automação de interface por visão computacional (imagem + OCR) para sistemas legados.
Author: Josias Azevedo da Silva
License: MIT License
        
        Copyright (c) 2026 Josias Azevedo da Silva
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
        ---
        
        This project extends and improves upon ideas and patterns from bot-vision-suite
        (https://github.com/matheuszwilk/bot-vision-suite), which is licensed under the
        MIT License by its respective authors. See that repository for the original
        copyright and license.
License-File: LICENSE
Keywords: automation,computer-vision,gui-automation,legacy-systems,ocr,opencv,oracle-forms,rpa,template-matching,tesseract
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.9
Requires-Dist: mss>=9.0
Requires-Dist: numpy>=1.21
Requires-Dist: opencv-python-headless>=4.5
Requires-Dist: pillow>=9.0
Requires-Dist: pyautogui>=0.9.53
Requires-Dist: pyperclip>=1.8
Requires-Dist: pytesseract>=0.3.8
Requires-Dist: pyyaml>=6.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Description-Content-Type: text/markdown

# pyvizion

Framework de **automação de interface (RPA desktop)** em Python: encontra botões e textos na tela, clica, digita e encadeia fluxos com recuperação de falhas.

![Python](https://img.shields.io/badge/python-3.9+-blue.svg)
![PyPI](https://img.shields.io/pypi/v/pyvizion?label=PyPI)
![Licença](https://img.shields.io/badge/licen%C3%A7a-MIT-green)

## O que é

O **pyvizion** é uma **extensão e evolução** da linha de automação desktop iniciada pelo [**bot-vision-suite**](https://pypi.org/project/bot-vision-suite/) ([repositório](https://github.com/matheuszwilk/bot-vision-suite)): mesma ideia central (imagem + OCR + backtrack), com melhorias de configuração, robustez em legado corporativo, CLI e manutenção contínua neste pacote (`pip install pyvizion`).

**Autoria:** desenvolvido e mantido por **Josias Azevedo da Silva**. Créditos ao **bot-vision-suite** e aos autores originais pela base conceitual e pela API de referência (MIT).

Foi feito para operar **sistemas legados no Windows** (e Linux/macOS onde aplicável): ERPs, Oracle Forms, Delphi/VB6, Java Swing, Citrix/RDP e terminais. Na prática você combina recortes PNG, OCR e teclas — sem depender de API do aplicativo.


| Camada | Tecnologia                                                                                           |
| ------ | ---------------------------------------------------------------------------------------------------- |
| Visão  | OpenCV — match multi-escala, DPI, vários monitores                                                   |
| Texto  | Tesseract — busca e leitura com perfis `fast` / `balanced` / `all`                                   |
| Ação   | PyAutoGUI — clique, digitação, atalhos                                                               |
| Fluxo  | **Backtrack** — se um passo falha, reexecuta o anterior e tenta de novo; sessões e listas de tarefas |


Instale com `pip install pyvizion`. Pacote: [pypi.org/project/pyvizion](https://pypi.org/project/pyvizion/). Licença **MIT**: [LICENSE](LICENSE).

```python
from pyvizion import Vizion

vz = Vizion()
vz.click_image("menu.png", backtrack=True)
vz.click_text("Relatórios", backtrack=True)
vz.click_text("Confirmar", sendtext="{enter}")
```

Versões antigas no PyPI (0.3.x) usavam outra API. Use a série **1.0+** (`from pyvizion import Vizion`).

---

## 📦 Instalação

```bash
pip install pyvizion
```

Uma instalação só, com tudo incluído e já no modo mais rápido.

**Tesseract** (só para funções de texto — `click_text`, `find_text`, `read_text`):

- **Windows:** [https://github.com/UB-Mannheim/tesseract/wiki](https://github.com/UB-Mannheim/tesseract/wiki) — instale em `C:\Program Files\Tesseract-OCR\` e marque **Portuguese**.
- **Linux:** `sudo apt-get install tesseract-ocr tesseract-ocr-por`
- **macOS:** `brew install tesseract tesseract-lang`

```bash
python -m pyvizion doctor      # confere dependências, Tesseract, idiomas e monitores
```

---

## 🎯 Uso rápido

### Backtrack entre métodos

```python
vz = Vizion()
vz.click_image("menu.png", backtrack=True)
vz.click_text("Relatórios", backtrack=True)   # falhou? reexecuta o anterior e tenta de novo
```

### Sessão de backtrack

```python
vz.start_task_session()
vz.click_image("button1.png", backtrack=True)
vz.click_text("Clientes", backtrack=True)
vz.click_text("Novo", backtrack=True)
successful, total = vz.end_task_session()
print(f"Sucesso: {successful}/{total}")
```

### Lista de tarefas

```python
from pyvizion import execute_tasks

tasks = [
    {"image": "button.png", "region": (100, 100, 200, 50), "confidence": 0.9,
     "specific": False, "backtrack": True, "delay": 1, "mouse_button": "left"},

    {"text": "Login", "region": (50, 50, 300, 100), "char_type": "letters",
     "backtrack": True, "sendtext": "usuario123{tab}senha{enter}"},

    {"type": "relative_image", "anchor_image": "warning_icon.png",
     "target_image": "ok_button.png", "max_distance": 200},

    {"type": "click", "x": 500, "y": 300, "mouse_button": "right"},
    {"type": "type_text", "text": "Hello World!"},
    {"type": "keyboard_command", "command": "Ctrl+S", "delay": 1},
]

execute_tasks(tasks)            # ou: Vizion().execute_tasks(tasks)
```

---

## 🔧 Métodos


| Método                                                                                                                                          | O que faz                                    |
| ----------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------- |
| `click_image(image_path, region, confidence, delay, mouse_button, max_attempts, backtrack, specific, sendtext, show_overlay)`                   | Encontra e clica numa imagem                 |
| `find_image(image_path, region, confidence, max_attempts, backtrack, specific, scales)`                                                         | Retorna `(x, y, largura, altura)` ou `None`  |
| `click_text(text, region, filter_type, delay, mouse_button, occurrence, backtrack, max_attempts, sendtext, confidence_threshold, show_overlay)` | Encontra e clica num texto                   |
| `find_text(text, region, filter_type, confidence_threshold, occurrence, max_attempts, backtrack)`                                               | Caixa do texto ou `None`                     |
| `click_relative_image(anchor_image, target_image, max_distance, confidence, target_region, delay, mouse_button, backtrack, max_attempts)`       | Clica no alvo mais próximo de uma âncora     |
| `find_relative_image(...)`                                                                                                                      | Caixa do alvo mais próximo da âncora         |
| `click_image_near_text(anchor_text, target_image, ...)`                                                                                         | Clica na imagem mais próxima de um texto     |
| `click_coordinates(x, y, delay, mouse_button, backtrack)` / `click_at(location, ...)`                                                           | Clique em coordenadas                        |
| `type_text(text, interval, delay, backtrack)`                                                                                                   | Digita texto (aceita macros)                 |
| `keyboard_command(command, delay, backtrack)`                                                                                                   | `"Ctrl+S"`, `"F7"`, `"Alt+Tab"`...           |
| `wait_for_image(...)` / `wait_for_text(...)` / `wait_until_gone(...)`                                                                           | Esperas                                      |
| `image_exists(...)` / `text_exists(...)` / `find_all_images(...)`                                                                               | Verificações                                 |
| `click_any([...])` / `find_any([...])`                                                                                                          | Primeiro alvo encontrado dentre alternativas |
| `read_text(region, single_line)`                                                                                                                | Lê o texto de uma área                       |
| `focus_window(title)` / `wait_window(title)` / `window_region(title)`                                                                           | Janelas                                      |
| `press(key, presses)` / `hotkey(*keys)` / `scroll(clicks)` / `drag(start, end)` / `screenshot(path)`                                            | Utilidades                                   |
| `execute_tasks(tasks)` / `execute_with_backtrack_between_tasks(tasks)`                                                                          | Listas de tarefas                            |
| `start_task_session()` / `end_task_session()`                                                                                                   | Sessão de backtrack                          |
| `configure_overlay(enabled, color, duration, width)` / `get_overlay_config()` / `test_overlay_colors()`                                         | Overlay visual                               |
| `list_windows()`                                                                                                                                | Títulos das janelas abertas                  |


**Parâmetros comuns**

- `region=(x, y, largura, altura)` — onde procurar. Use sempre que puder (OCR ~6× mais rápido, sem falsos positivos). `vz.window_region("Título")` devolve a área de uma janela.
- `specific=True` — busca só na `region`. `specific=False` — tenta a região e depois a tela inteira, em escalas 0,8× a 1,2×.
- `mouse_button` — `"left"`, `"right"`, `"double"`, `"move_to"` (só passa o mouse), `"middle"`, `"triple"`.
- `filter_type` — `"letters"`, `"numbers"` ou `"both"`. `occurrence=2` — a segunda ocorrência em ordem de leitura.
- `offset=(dx, dy)` — clica deslocado do alvo (ex.: o campo à direita do rótulo).
- `timeout=` — espera o alvo aparecer por até N segundos.

Os métodos também existem como funções, sem instância: `from pyvizion import click_image, find_text, execute_tasks`.

---

## 📝 Preenchendo campos (`sendtext`)

`sendtext` é o texto digitado **logo depois do clique**. Dentro dele, palavras entre chaves `{ }` são **teclas**:

```python
vz.click_text("Usuário", offset=(150, 0), sendtext="admin{tab}senha123{enter}")
```

O que acontece, em ordem:


| #   | Trecho     | O pyvizion faz                                               |
| --- | ---------- | ------------------------------------------------------------ |
| 1   | *(clique)* | clica 150 px à direita do rótulo "Usuário" (dentro do campo) |
| 2   | `admin`    | digita "admin"                                               |
| 3   | `{tab}`    | aperta **Tab** → vai para o campo de senha                   |
| 4   | `senha123` | digita "senha123"                                            |
| 5   | `{enter}`  | aperta **Enter** → confirma                                  |


### Receitas do dia a dia


| Quero...                              | `sendtext`                    |
| ------------------------------------- | ----------------------------- |
| digitar num campo vazio               | `"12345"`                     |
| **substituir** o que já está no campo | `"{ctrl}a{del}12345"`         |
| digitar e ir para o próximo campo     | `"12345{tab}"`                |
| digitar e confirmar                   | `"12345{enter}"`              |
| preencher dois campos seguidos        | `"01/01/2026{tab}31/01/2026"` |
| pular campos                          | `"{tab*3}"`                   |
| esperar o sistema reagir              | `"12345{tab}{wait 1}"`        |


Prefira **um passo por campo** — fica mais fácil de ler e, se algo falhar, o log mostra exatamente onde:

```python
vz.click_text("Data Inicial", offset=(150, 0), sendtext="{ctrl}a{del}01/01/2026")
vz.click_text("Data Final",   offset=(150, 0), sendtext="{ctrl}a{del}31/01/2026")
vz.keyboard_command("Enter")
```

### Teclas disponíveis


| Macro                               | Tecla                         |
| ----------------------------------- | ----------------------------- |
| `{enter}` `{tab}` `{esc}` `{space}` | Enter, Tab, Esc, Espaço       |
| `{del}` `{backspace}` `{insert}`    | Delete, Backspace, Insert     |
| `{up}` `{down}` `{left}` `{right}`  | setas                         |
| `{home}` `{end}` `{pgup}` `{pgdn}`  | Home, End, Page Up, Page Down |
| `{f1}` … `{f12}`                    | teclas de função              |
| `{ctrl}a` `{alt}f` `{shift}x`       | modificador + a próxima letra |
| `{ctrl+shift+s}` `{alt+f4}`         | combinação completa           |
| `{tab*3}` `{down 5}`                | repetição                     |
| `{wait 1.5}` `{wait 300ms}`         | pausa                         |
| `{{` `}}`                           | as chaves `{` e `}` literais  |


- Maiúsculas não importam: `{Enter}` = `{enter}`.
- Chaves com algo que não é tecla são digitadas normalmente: `"valor {total}"` digita exatamente isso.
- Acentos e `ç` saem corretos, e a área de transferência do usuário é restaurada.
- Em **Citrix/RDP/terminais** que não aceitam colar, use `Vizion({"typing_mode": "type"})`.

Para apertar **só uma tecla**, sem texto: `vz.keyboard_command("Ctrl+S")`, `vz.press("tab", 3)`, `vz.hotkey("ctrl", "shift", "s")`.

---

## 📋 Tarefas: todos os tipos

```python
tasks = [
    {"type": "focus_window", "title": "Oracle Applications"},
    {"type": "wait_image", "image": "tela_principal.png", "timeout": 30},
    {"text": "Cliente", "offset": (160, 0), "sendtext": "12345{enter}", "required": True},
    {"type": "wait_image", "image": "ampulheta.png", "gone": True, "timeout": 60},
    {"type": "relative_image", "anchor_text": "Pedido", "target_image": "lupa.png"},
    {"image": "popup_aviso.png", "optional": True, "sendtext": "{enter}"},
    {"type": "wait_text", "text": "Registro salvo", "timeout": 15},
    {"type": "wait", "seconds": 1},
    {"type": "scroll", "clicks": -5},
]
```


| Tipo                               | Chaves                                                                           |
| ---------------------------------- | -------------------------------------------------------------------------------- |
| imagem                             | `image`, `region`, `confidence`, `specific`, `scales`                            |
| texto                              | `text`, `region`, `char_type`, `occurrence`, `confidence_threshold`              |
| `relative_image`                   | `anchor_image` ou `anchor_text`, `target_image`, `max_distance`, `target_region` |
| `click`                            | `x`, `y`                                                                         |
| `type_text`                        | `text`, `interval`                                                               |
| `keyboard_command`                 | `command`                                                                        |
| `wait_image` / `wait_text`         | `image`/`text`, `timeout`, `gone`                                                |
| `wait` · `focus_window` · `scroll` | `seconds` · `title` · `clicks`, `x`, `y`                                         |


Chaves de qualquer tarefa: `mouse_button`, `delay`, `sendtext`, `offset`, `backtrack`, `max_attempts`, `timeout`, `optional` (falha não conta nem faz backtrack), `required` (falha interrompe a lista), `show_overlay`, `click_hold`.

---

## ⚙️ Configuração

```python
vz = Vizion({
    "confidence_threshold": 80.0,
    "tesseract_lang": "por",
    "show_overlay": False,
    "image_folders": ["./imagens"],
    "save_failure_screenshots": True,
})

vz.config.set("show_overlay", True)          # alterar depois
vz = Vizion("pyvizion.json")                  # ou de um arquivo .json/.yaml
```


| Chave                                                                   | Padrão                              | Descrição                                            |
| ----------------------------------------------------------------------- | ----------------------------------- | ---------------------------------------------------- |
| `confidence_threshold`                                                  | `75.0`                              | Limiar do OCR (0–100)                                |
| `default_confidence`                                                    | `0.9`                               | Confiança padrão de imagens nas tarefas              |
| `tesseract_path` / `tessdata_path`                                      | auto                                | Caminhos do Tesseract                                |
| `tesseract_lang`                                                        | auto (`por` se instalado)           | Idioma(s): `"por"`, `"por+eng"`                      |
| `image_processing_methods`                                              | `"all"`                             | `"all"`, `"balanced"`, `"fast"` ou lista de técnicas |
| `ocr_large_image_methods`                                               | `"fast"`                            | Técnicas usadas na tela inteira                      |
| `preprocessing_enabled`                                                 | `True`                              | `False` = OCR só na imagem original                  |
| `ocr_upscale` / `ocr_workers`                                           | `2.0` / auto                        | Ampliação de áreas pequenas / paralelismo            |
| `ocr_fuzzy` / `ocr_fuzzy_threshold`                                     | `True` / `0.8`                      | Tolerância a erros do OCR                            |
| `grayscale` · `scales` · `dpi_scales` · `edge_fallback`                 | `True` · 0,8–1,2 · `True` · `False` | Busca de imagem                                      |
| `min_confidence` · `retry_delay`                                        | `0.7` · `0.5`                       | Tentativas                                           |
| `overlay_enabled`                                                       | `True`                              | Liga/desliga o sistema de overlay                    |
| `image_folders`                                                         | `[]`                                | Pastas onde procurar imagens                         |
| `typing_mode` · `restore_clipboard` · `typing_interval`                 | `"paste"` · `True` · `0.02`         | Digitação                                            |
| `click_hold` · `move_pause` · `movement_duration`                       | `0` · `0.15` · `0.1`                | Ritmo do mouse                                       |
| `failsafe`                                                              | `True`                              | Mouse no canto superior esquerdo interrompe          |
| `show_overlay` · `overlay_color` · `overlay_duration` · `overlay_width` | `False` · `red` · `1000` · `4`      | Retângulo sobre o alvo antes do clique (depuração)   |
| `save_failure_screenshots` · `failure_screenshot_dir`                   | `False` · `pyvizion_failures`       | Print a cada falha                                   |
| `stop_on_failure` · `max_backtrack_attempts`                            | `False` · `2`                       | Listas de tarefas                                    |
| `log_level`                                                             | `"INFO"`                            |                                                      |


---

## 🛡️ Robustez para sistemas legados

- **DPI e vários monitores** — modo DPI por monitor, captura de todos os monitores (inclusive coordenadas negativas); a escala de cada monitor (125%, 150%) entra automaticamente na busca.
- **OCR tolerante** — ignora acentos e maiúsculas, corrige `0/O`, `1/l`, `5/S`, aceita pequenas diferenças e palavras coladas/quebradas, amplia áreas pequenas, inverte texto claro em fundo escuro, roda em paralelo e usa consenso entre leituras.
- **Captura rápida** — ~5 ms por área.
- **Esperas**, **alternativas** (`click_any`), **âncoras** (`click_relative_image`, `click_image_near_text`), **clique deslocado** (`offset`), **leitura de campos** (`read_text`), **foco de janela**.
- **Listas seguras** — `required`, `optional`, `stop_on_failure`.
- **Diagnóstico** — prints de falha e `python -m pyvizion doctor`.

---

## 🛠️ Linha de comando

```bash
python -m pyvizion doctor               # ambiente
python -m pyvizion screenshot tela.png  # print para recortar suas imagens
python -m pyvizion position             # x/y do mouse em tempo real (Ctrl+C)
python -m pyvizion pick                 # marca região (2 cantos), OCR e trechos para colar no código
python -m pyvizion ocr 100 200 300 40   # lê o texto de uma área (se já souber x,y,w,h)
```

---

## 💡 Dicas

1. Use **`region`** (ou `vz.window_region("Título")`) sempre que puder.
2. Recorte **imagens pequenas e únicas** (o ícone, não o botão inteiro), em PNG.
3. Prefira **`wait_until_gone("ampulheta.png")`** a `delay` fixo.
4. Comece com **`focus_window("Título")`**.
5. Instale o idioma **`por`** do Tesseract.
6. Para parar o robô, leve o mouse ao **canto superior esquerdo** da tela.

---

## 🧪 Desenvolvimento

```bash
pip install -e .[dev]
pytest
```

## 📄 Licença e créditos

Licenciado sob **MIT** — texto completo em [LICENSE](LICENSE).

| | |
|---|---|
| **pyvizion** | Copyright © 2026 Josias Azevedo da Silva |
| **Base** | [bot-vision-suite](https://github.com/matheuszwilk/bot-vision-suite) (MIT) — projeto de referência para automação GUI com OCR e detecção de imagens |

O pyvizion **não substitui** o bot-vision-suite no PyPI; é um pacote separado com melhorias e API `Vizion` / `from pyvizion import …`, mantido por Josias Azevedo da Silva.