Metadata-Version: 2.4
Name: pickerxl
Version: 0.2.5
Summary: PickerXL is a large deep learning model to measure arrival times from noisy seismic signals.
Author-email: "Chengping Chai, Derek Rose, Scott Stewart, Nathan Martindale, Mark Adams, Lisa Linville, Christopher Stanley, Anibely Torres Polanco and Philip Bingham" <chaic@ornl.gov>
License-Expression: GPL-3.0-or-later
Project-URL: Homepage, https://github.com/ornl/picker-xl
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch>=2.2.0
Requires-Dist: numpy>=1.26.4
Requires-Dist: pytorch-lightning>=2.2.4
Requires-Dist: obspy>=1.4.1
Provides-Extra: development
Requires-Dist: flake8; extra == "development"
Requires-Dist: black; extra == "development"
Requires-Dist: pre-commit; extra == "development"
Provides-Extra: examples
Requires-Dist: h5py>=3.11.0; extra == "examples"
Requires-Dist: matplotlib>=3.8.0; extra == "examples"
Dynamic: license-file

# PickerXL

A large deep learning model to measure arrival times from noisy seismic signals. The model was developed by Chengping Chai, Derek Rose, Scott Stewart, Nathan Martindale, Mark Adams, Lisa Linville, Christopher Stanley, Anibely Torres Polanco and Philip Bingham.


## Introduction

This model is trained on STEAD (Mousavi et al., 2019) for Primary (P) and Secondary (S) wave arrival picking. The model was trained using earthquake data at local distances (0-350 km). The model uses 57 s long three-component seismograms sampled at 100 Hz as input. The required channel order is East, North, and Vertical. The output of the model are probability channels corresponding to P wave, S wave, and noise. These probabilities can be used to compute P- and S-wave arrival times.

`Picker.predict_probability()` returns a NumPy array with shape `(batch, 3, n_samples)`. If an input trace is longer than `5700` samples, PickerXL segments it internally, applies the model window by window, and stitches the probabilities back to the original sample length.

`Picker.predict_arrivals()` applies `obspy.signal.trigger.trigger_onset()` to the P and S probability channels. It returns NaN-padded arrays of trigger picks, where each row corresponds to one input sample and each column corresponds to one detected trigger.

## Installation

Creating a virtual environment is highly encouraged. 

```
pip install pickerxl
```

To run the bundled HDF5 and plotting examples, install the optional example dependencies too:

```
pip install "pickerxl[examples]"
```

To install only for CPU (e.g., macOS), you may use the following commands.

```
pip3 install torch torchvision torchaudio
pip3 install pickerxl
```


## Example Usage

```python
from pickerxl.pickerxl import Picker
import numpy as np
import h5py
model = Picker()
fid = h5py.File("example_waveforms.h5", "r")
data_group = fid["data"]
example_data = []
true_p_index = []
true_s_index = []
for akey in data_group.keys():
    dataset = data_group[akey]
    example_data.append(dataset[...])
    true_p_index.append(float(dataset.attrs["p_arrival_sample"]))
    true_s_index.append(float(dataset.attrs["s_arrival_sample"]))
fid.close()
preds = model.predict_probability(example_data)
p_index, s_index = model.predict_arrivals(example_data)
# `preds` has shape (batch, 3, n_samples).
# Input channels must be ordered as East, North, Vertical.
# `p_index` and `s_index` are NaN-padded arrays of trigger picks.
print("True P-wave arrival index:", true_p_index)
print("Predicted P-wave arrival index:", p_index)
print("True S-wave arrival index:", true_s_index)
print("Predicted S-wave arrival index:", s_index)
```

To return trigger peak probabilities together with the trigger indices:

```python
p_index, s_index, p_prob, s_prob = model.predict_arrivals(
    example_data,
    return_probabilities=True,
)
```

## MiniSEED Example

The repository includes a MiniSEED example at [tests/example_waveform.mseed](/Users/c09/Mycodes/SteelThread/picker-xl/tests/example_waveform.mseed). The test script preprocesses the stream before applying PickerXL by:

1. merging traces and interpolating gaps,
2. removing the mean,
3. removing the linear trend,
4. resampling to `100 Hz`,
5. applying a `1-45 Hz` bandpass filter.

The waveform passed to PickerXL must be stacked in East, North, Vertical order. In the test script, MiniSEED traces are explicitly sorted by `E`, `N`, and `Z` before stacking.

Example:

```python
from obspy import read
import numpy as np
from pickerxl.pickerxl import Picker

stream = read("tests/example_waveform.mseed")
stream.merge(method=1, fill_value="interpolate")
stream.detrend("demean")
stream.detrend("linear")
stream.resample(100.0)
stream.filter("bandpass", freqmin=1.0, freqmax=45.0, corners=4, zerophase=True)

channel_order = ("E", "N", "Z")
component_map = {}
sorted_traces = sorted(
    stream,
    key=lambda trace: channel_order.index(trace.stats.channel[-1].upper())
    if trace.stats.channel[-1].upper() in channel_order
    else len(channel_order),
)
for trace in sorted_traces:
    component = trace.stats.channel[-1].upper()
    if component in channel_order and component not in component_map:
        component_map[component] = trace

ordered_components = tuple(sorted(component_map, key=channel_order.index))
npts = min(component_map[component].stats.npts for component in ordered_components)
waveform = np.stack(
    [
        component_map[component].data[:npts].astype(np.float32)
        for component in ordered_components
    ]
)
# The stacked waveform is ordered as East, North, Vertical.

model = Picker()
preds = model.predict_probability(waveform)
p_index, s_index = model.predict_arrivals(waveform)
print("MiniSEED prediction shape:", preds.shape)
print("MiniSEED P-wave trigger indices:", p_index[0])
print("MiniSEED S-wave trigger indices:", s_index[0])
```

## Test the Package

First, go to the top directory of the package. Then run:

```
cd tests
python run_tests.py
```

![example image](images/example_waveform_2.png)

The test script also reads `tests/example_waveform.mseed`, applies the preprocessing steps above, and saves a MiniSEED plot to `tests/example_waveform_mseed.jpg`.


## Known Limitations

* The model may have a less-than-optimal performance for earthquake data outside of a source-receiver distance range of 10–110 km and a magnitude range of 0–4.5 because of biases in the training data.
* The model may produce false detections when applied to continuous seismic data.
* The model may not perform well for earthquake data at larger distances or for non-earthquake sources.

## Reference

Chengping Chai, Derek Rose, Scott Stewart, Nathan Martindale, Mark Adams, Lisa Linville, Christopher Stanley, Anibely Torres Polanco, Philip Bingham; PickerXL, A Large Deep Learning Model to Measure Arrival Times from Noisy Seismic Signals. Seismological Research Letters 2025; 96 (4):2394-2404. doi: https://doi.org/10.1785/0220240353

## License

GNU GENERAL PUBLIC LICENSE version 3

## Credit

The architecture of the deep learning model was adapted from SeisBench (Woollam et al., 2022).
