Metadata-Version: 2.4
Name: eubi_bridge
Version: 0.1.3
Summary: A package for converting datasets to OME-Zarr format.
Home-page: https://github.com/Euro-BioImaging/EuBI-Bridge
Author: Bugra Özdemir
Author-email: Bugra Ã–zdemir <bugraa.ozdemir@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/Euro-BioImaging/EuBI-Bridge
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.11,<3.13
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: aicspylibczi>=0.0.0
Requires-Dist: asciitree>=0.3.3
Requires-Dist: bfio>=0.0.0
Requires-Dist: bioformats_jar>=0.0.0
Requires-Dist: bioio-base==3.3.0
Requires-Dist: bioio-bioformats==1.1.0
Requires-Dist: bioio-czi==2.1.0
Requires-Dist: bioio-imageio==1.1.0
Requires-Dist: bioio-lif==1.1.0
Requires-Dist: bioio-nd2==1.1.0
Requires-Dist: bioio-ome-tiff-fork-by-bugra==0.0.1b2
Requires-Dist: bioio-tifffile-fork-by-bugra>=0.0.1b2
Requires-Dist: cmake==4.0.2
Requires-Dist: dask>=2024.12.1
Requires-Dist: dask-jobqueue>=0.0.0
Requires-Dist: distributed>=2024.12.1
Requires-Dist: elementpath==5.0.1
Requires-Dist: fasteners==0.19
Requires-Dist: imageio==2.27.0
Requires-Dist: imageio-ffmpeg==0.6.0
Requires-Dist: install-jdk
Requires-Dist: lz4>=4.4.4
Requires-Dist: matplotlib
Requires-Dist: micro-reader>=0.0.1
Requires-Dist: natsort>=0.0.0
Requires-Dist: nd2>=0.0.0
Requires-Dist: numpy>=1.24
Requires-Dist: pydantic>=2.11.7
Requires-Dist: pylibczirw>=0.0.0
Requires-Dist: PyQt6<=6.9.1,>=6.6.0
Requires-Dist: pytest>=7.0
Requires-Dist: pytest-cov>=4.0
Requires-Dist: readlif==0.6.5
Requires-Dist: s3fs>=0.0.0
Requires-Dist: scipy>=1.8
Requires-Dist: tensorstore>=0.0.0
Requires-Dist: tifffile<2026,>=2025.5.21
Requires-Dist: validators==0.35.0
Requires-Dist: xarray>=0.0.0
Requires-Dist: xmlschema>=0.0.0
Requires-Dist: xmltodict==0.14.2
Requires-Dist: zarr<3.2,>=3.0
Requires-Dist: zstandard>=0.0.0
Requires-Dist: blosc2>=3.7.1
Requires-Dist: aiofiles>=24.1.0
Requires-Dist: psutil>=7.0.0
Requires-Dist: rich>=14.1.0
Requires-Dist: h5py
Requires-Dist: fire>=0.0.0
Provides-Extra: flow
Requires-Dist: eubi-flow>=0.1.2b8; extra == "flow"
Provides-Extra: annotate
Requires-Dist: eubi-annotate>=0.1.2b8; extra == "annotate"
Provides-Extra: deep
Requires-Dist: eubi-deep>=0.1.2b8; extra == "deep"
Provides-Extra: docs
Requires-Dist: mkdocs-material>=9.5; extra == "docs"
Requires-Dist: pymdown-extensions>=10.0; extra == "docs"
Dynamic: author
Dynamic: home-page
Dynamic: license-file
Dynamic: requires-python

# EuBI-Bridge  

[![Documentation](https://img.shields.io/badge/documentation-online-green)](https://euro-bioimaging.github.io/EuBI-Bridge/)

EuBI-Bridge is a tool for **distributed conversion of microscopy image collections to OME-Zarr**. Conversion can be performed through either a command-line interface (CLI) or a graphical user interface (GUI). The installation described below provides access to both interfaces.

A key feature of EuBI-Bridge is **aggregative conversion**, which combines multiple input images along user-specified axes into a single OME-Zarr container. This is particularly useful when images belong to the same multidimensional dataset but are stored as separate, unlinked files. By aggregating these files during conversion, EuBI-Bridge can reconstruct the intended multidimensional dataset.

EuBI-Bridge is built on several powerful libraries, including `zarr`, `tensorstore` and `dask`, among others.

Input files are read with [micro-reader](https://pypi.org/project/micro-reader/) by default. It reads the common microscopy formats without Java. If a file cannot be read by `micro-reader`, EuBI-Bridge falls back to other readers (such as `bioio`- and `Bio-Formats`-based readers) for that particular file. Java only starts when this happens. 


## Installation

**Important: EuBI-Bridge is currently only compatible with Python 3.11 or 3.12 due to conflicting dependencies.
We are working on supporting a wider range of Python versions in future releases.**

### Recommended: conda

The recommended way to install EuBI-Bridge is in a conda environment. This
route also installs Qt, so the graphical interface works without any further
setup:

```bash
mamba create -n eubizarr openjdk=11.* maven python=3.12 pyqt6=6.8.1
```

Then install EuBI-Bridge via pip in the conda environment:

```bash
conda activate eubizarr
pip install --no-cache-dir "eubi-bridge==0.1.3"
eubi reset_config   # see "After installing or upgrading" below
```

### Alternative: pip only

Installing into a plain virtual environment also works, provided your Python is
version 3.11 or 3.12:

```bash
python -m venv venv # Python must be either version 3.11 or 3.12.
source venv/bin/activate
pip install "eubi-bridge==0.1.3" # installs both GUI and CLI
eubi reset_config   # see "After installing or upgrading" below
```

**On Linux this installs the graphical interface but not the Qt system
libraries it needs at runtime**, so the graphical interface may fail to start. See
[Linux: the GUI does not start](#linux-the-gui-does-not-start) below for the fix.
The command-line interface is unaffected.

### After installing or upgrading

Run this once:

```bash
eubi reset_config
```

EuBI-Bridge keeps your settings in a configuration file. That file is written the
first time you run the tool, and later versions do not change it. So if you have
used EuBI-Bridge before, you keep the old defaults until you reset it.

- Resetting gives you the current defaults. For example, files are now read with
  micro-reader, and Java only starts if a file needs Bio-Formats.
- It overwrites your default settings. Run `eubi show_config` first if you
  want to note them down. Configurations you saved under a name are kept.
- On a fresh install it does no harm.

### Troubleshooting

#### Linux: the GUI does not start

This applies to the pip-only route. On Linux, `pip` installs PyQt6 but not the
system libraries it needs at runtime, so the `eubi-gui` command (which is run to launch the GUI) may stop with an error such
as:

```bash
EuBI-Bridge could not start its graphical interface: libEGL.so.1: cannot open shared object file
```

Install the Qt system libraries with your package manager:

```bash
# Debian / Ubuntu
sudo apt install libegl1 libgl1 libxkbcommon-x11-0 libdbus-1-3 libxcb-cursor0

# Fedora / RHEL
sudo dnf install mesa-libEGL mesa-libGL libxkbcommon-x11 dbus-libs xcb-util-cursor
```

If you cannot install system packages, use the recommended conda route instead:
`mamba create ... pyqt6=6.8.1` brings its own Qt libraries and avoids this
problem entirely. The `eubi` command-line interface needs none of them and
works either way.

#### Building wheel errors

If you receive a `Building wheel` error such as:

```bash
  Building wheel for ... error
  error: subprocess-exited-with-error
  
  × python setup.py bdist_wheel did not run successfully.
  │ exit code: 1
```
then try the following:

```bash
# In the `eubizarr` environment
mamba install cmake zlib boost # preinstall dependencies that can help build from source
pip install --no-cache-dir "eubi-bridge==0.1.3" # try installing again with the dependencies available
eubi reset_config
```

## Documentation

Find the documentation for EuBI-Bridge [here](https://euro-bioimaging.github.io/EuBI-Bridge/)

## Graphical interface

Launch the graphical interface by running `eubi-gui` in the terminal. The four short videos below walk you through one complete job: converting several images to OME-Zarr, and then inspecting one of the outputs.

The videos cover the common scenario: one-to-one conversion from a collection of files. More advanced features such as aggregative conversions (concatenation of multiple files), editable batch tables, custom configuration files and editing output metadata are all supported by the interface but not yet demonstrated in the videos; demos for those will be provided soon. In the meantime the [documentation](https://euro-bioimaging.github.io/EuBI-Bridge/) describes the API, and every control in the interface carries a tooltip. 

### 1. Selecting input and output folders

Navigate to your images, filter them with include and exclude patterns (which accept star `*` expressions), tick the files to be converted, and choose where the OME-Zarr output should be written.

https://github.com/user-attachments/assets/0a474606-d443-40f0-8f36-23053c61cc81

### 2. Choosing conversion parameters

Work through the parameter tabs. The tabs cover different parameter categories ranging from resource allocation, through reader options, chunking, sharding and compression, downscaling, channel and pixel metadata overrides, and more. Each field explains itself on hover, and settings can be saved to a configuration file for reuse.

https://github.com/user-attachments/assets/c7e3ab3b-37cf-4565-a37a-f2b6fa985961

### 3. Running a conversion

Start the job and watch it progress. The log reports each file as it is read, converted and downscaled, so a failure points at the file that caused it.

https://github.com/user-attachments/assets/3492e69d-ea83-4363-b3c7-90a9c00be76c

### 4. Inspecting the result

Switch to the Inspect tab to open one of the converted OME-Zarr stores from the list in the sidebar. In the Inspect/Metadata path, find full OME-Zarr metadata, including axis names, units, pixel sizes, the chunk and shard layout and compression per pyramid level. Pixel sizes and units can be corrected here and saved back into the OME-Zarr. Switch to the Inspect/Viewer path and visualise the OME-Zarr. Adjust a specific resolution layer via Zoom slider. Browse through the z slices or timepoints. Pan the frame of view using the left mouse button. Add or remove specific channels, and adjust the channel metadata such as colours, viewed pixel range and channel names. The updated visualisation metadata can also be saved back into the OME-Zarr.

https://github.com/user-attachments/assets/08149eb8-1cb6-403b-81a9-668ce4637f24

The same recordings are attached to the
[v0.1.2-media release](https://github.com/Euro-BioImaging/EuBI-Bridge/releases/tag/v0.1.2-media) for viewing outside GitHub. They are not part of
the repository, so cloning or pip-installing the package does not download them.

## Command line interface  

### Unary Conversion  

Given a dataset structured as follows: 

```bash
multichannel_timeseries
├── Channel1-T0001.tif
├── Channel1-T0002.tif
├── Channel1-T0003.tif
├── Channel1-T0004.tif
├── Channel2-T0001.tif
├── Channel2-T0002.tif
├── Channel2-T0003.tif
└── Channel2-T0004.tif
```  

To convert each TIFF into a separate OME-Zarr container (unary conversion):  

```bash
eubi to_zarr multichannel_timeseries multichannel_timeseries_zarr
```  

By default this writes OME-Zarr 0.4 (zarr version 2). Use `--ome_zarr_version`
to choose the version. For OME-Zarr 0.5 (zarr version 3, which supports sharding):

```bash
eubi to_zarr multichannel_timeseries multichannel_timeseries_zarr --ome_zarr_version 0.5
```  

`--zarr_format` still works, but it is deprecated. Use `--ome_zarr_version` instead.

Both of these commands will perform unary conversion, resulting in the following output:  

```bash
multichannel_timeseries_zarr
├── Channel1-T0001.zarr
├── Channel1-T0002.zarr
├── Channel1-T0003.zarr
├── Channel1-T0004.zarr
├── Channel2-T0001.zarr
├── Channel2-T0002.zarr
├── Channel2-T0003.zarr
└── Channel2-T0004.zarr
```  

Use **wildcards** to specifically convert the images belonging to Channel1:

```bash
eubi to_zarr "multichannel_timeseries/Channel1*" multichannel_timeseries_channel1_zarr
```

### Aggregative Conversion (Concatenation Along Dimensions)  

To concatenate images along specific dimensions, EuBI-Bridge needs to be informed
of file patterns that specify image dimensions. For this example,
the file pattern for the channel dimension is `Channel`, which is followed by the channel index,
and the file pattern for the time dimension is `T`, which is followed by the time index.

To concatenate along the **time** dimension:

```bash
eubi to_zarr multichannel_timeseries multichannel_timeseries_concat_zarr \
--channel_tag Channel \
--time_tag T \
--concatenation_axes t
```  

Output:  

```bash
multichannel_timeseries_concat_zarr
├── Channel1-T_tset.zarr
└── Channel2-T_tset.zarr
```  

**Important note:** if the `--channel_tag` was not provided, the tool would not be aware
of the multiple channels in the image and try to concatenate all images into a single one-channeled OME-Zarr. Therefore, 
when an aggregative conversion is performed, all dimensions existing in the input files must be specified via their respective tags. 

For multidimensional concatenation (**channel** + **time**):

```bash
eubi to_zarr multichannel_timeseries multichannel_timeseries_concat_zarr \
--channel_tag Channel \
--time_tag T \
--concatenation_axes ct
```  

Note that both axes are specified via the argument `--concatenation_axes ct`.

Output:

```bash
multichannel_timeseries_concat_zarr
└── Channel_cset-T_tset.zarr
```  

### Handling Nested Directories  

For datasets stored in nested directories such as:  

```bash
multichannel_timeseries_nested
├── Channel1
│   ├── T0001.tif
│   ├── T0002.tif
│   ├── T0003.tif
│   ├── T0004.tif
├── Channel2
│   ├── T0001.tif
│   ├── T0002.tif
│   ├── T0003.tif
│   ├── T0004.tif
```  

EuBI-Bridge automatically detects the nested structure. To concatenate along both channel and time dimensions:  

```bash
eubi to_zarr \
multichannel_timeseries_nested \
multichannel_timeseries_nested_concat_zarr \
--channel_tag Channel \
--time_tag T \
--concatenation_axes ct
```  

Output:  

```bash
multichannel_timeseries_nested_concat_zarr
└── Channel_cset-T_tset.zarr
```  

To concatenate along the channel dimension only:  

```bash
eubi to_zarr \
multichannel_timeseries_nested \
multichannel_timeseries_nested_concat_zarr \
--channel_tag Channel \
--time_tag T \
--concatenation_axes c
```  

Output:  

```bash
multichannel_timeseries_nested_concat_zarr
├── Channel_cset-T0001.zarr
├── Channel_cset-T0002.zarr
├── Channel_cset-T0003.zarr
└── Channel_cset-T0004.zarr
```  

### Selective Data Conversion    

To recursively select specific files for conversion, wildcard patterns can be used. 
For example, to concatenate only **timepoint 3** along the channel dimension:  

```bash
eubi to_zarr \
"multichannel_timeseries_nested/**/*T0003*" \
multichannel_timeseries_nested_concat_zarr \
--channel_tag Channel \
--time_tag T \
--concatenation_axes c
```  

Output:  

```bash
multichannel_timeseries_nested_concat_zarr
└── Channel_cset-T0003.zarr
```  

**Note:** When using wildcards, the input directory path must be enclosed 
in quotes as shown in the example above.  

### Handling Categorical Dimension Patterns  

For datasets where channel names are categorical such as in:

```bash
blueredchannels_timeseries
├── Blue-T0001.tif
├── Blue-T0002.tif
├── Blue-T0003.tif
├── Blue-T0004.tif
├── Red-T0001.tif
├── Red-T0002.tif
├── Red-T0003.tif
└── Red-T0004.tif
```

Specify categorical names as a comma-separated list:  

```bash
eubi to_zarr \
blueredchannels_timeseries \
blueredchannels_timeseries_concat_zarr \
--channel_tag Blue,Red \
--time_tag T \
--concatenation_axes ct
```  

Output:  

```bash
blueredchannels_timeseries_concat_zarr
└── BlueRed_cset-T_tset.zarr
```  

Note that the categorical names are aggregated in the output OME-Zarr name.  


With nested input structure such as in:  

```bash
blueredchannels_timeseries_nested
├── Blue
│   ├── T0001.tif
│   ├── T0002.tif
│   ├── T0003.tif
│   └── T0004.tif
└── Red
    ├── T0001.tif
    ├── T0002.tif
    ├── T0003.tif
    └── T0004.tif
```  

One can run the exact same command:

```bash
eubi to_zarr \
blueredchannels_timeseries_nested \
blueredchannels_timeseries_nested_concat_zarr \
--channel_tag Blue,Red \
--time_tag T \
--concatenation_axes ct
```  

Output:  

```bash
blueredchannels_timeseries_nested_concat_zarr
└── BlueRed_cset-T_tset.zarr
```

## Additional Notes

- EuBI-Bridge is in the **beta stage**, and significant updates may be expected.
- **Community support:** Questions and contributions are welcome! Please report any issues.


