Metadata-Version: 2.4
Name: picblocks
Version: 2.1.0
Summary: A library for code similarity estimation using PIC hashing over basic blocks.
Author-email: Daniel Plohmann <daniel.plohmann@mailbox.org>
License-Expression: BSD-2-Clause
Project-URL: Homepage, https://github.com/danielplohmann/picblocks
Classifier: Development Status :: 4 - Beta
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Security
Classifier: Topic :: Software Development :: Disassemblers
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: smda>=4.2.13
Provides-Extra: web
Requires-Dist: flask; extra == "web"
Requires-Dist: werkzeug; extra == "web"
Requires-Dist: waitress; extra == "web"
Requires-Dist: pymongo; extra == "web"
Requires-Dist: tqdm; extra == "web"
Provides-Extra: dev
Requires-Dist: build; extra == "dev"
Requires-Dist: pytest; extra == "dev"
Requires-Dist: requests; extra == "dev"
Requires-Dist: ruff; extra == "dev"
Requires-Dist: twine; extra == "dev"
Requires-Dist: ty; extra == "dev"
Dynamic: license-file

# PicBlocks

An experimental project using position-independent code hashing over basic blocks for code similarity estimation.

## Usage

Both module files in `./picblocks` and in `./utils` are runnable and contain examples of their usage:

* `$ python -m picblocks.blockhasher <target_binary_path>` - produces a `block-report` for a single binary.
* `$ python -m picblocks.blockhashmatcher <block_reports_path>` - creates a new `./db/picblocksdb.json` from the `block-reports` located in `<block_reports_path>`
* `$ python -m picblocks.blockhashmatcher <block_reports_path> <target_binary_path>` - matches a binary against data stored in `./db/picblocksdb.json` if it exists, or otherwise creates `./db/picblocksdb.json` from the `block-reports` located in `<block_reports_path>`
* `$ python -m utils.import_picblocksdb_to_mongo` assumes some mongodb configurations (please check inside the file to adapt to yours) it merely takes the json generated DB into a most easy to manage (and query)  mongodb. 
* `$ python -m utils.make_stats` it assumes a mongodb connection (please check inside the file to adapt to yours), the generated json db into `db/picblocksdb.json` (you can change it directly in the relative varible) and the generated blocks report into `./block-reports/` folder. It builds up some statistics about detections and DB composition. The results would be available in a dedicated (and very simple) stats web ui. Without MongoDB it writes `db/stats.json` instead. 

## Creating a Database

The script `hash_malpedia.py` is an example of how to process a collection of binaries into `./block-reports`, which will then be aggreated into a `./db/picblocksdb.json`.

## Database Evaulation

In oder to quantify and to measure the quality of your detection rate you should check some basic informations about tests run against your db. 
The simple (and preliminary) script named `make_stats.py` would build up some initial stats for you about detection rates. 
It assumes to have a mongodb connection, the generated json db into `db/picblocksdb.json` (you can change it directly in the relative varible) and the generated blocks reports into `./block-reports/` folder (you can change it directly on the specified variable). 
Once you run it, it takes every single block report and check it against the json database.
It build some stats and saves all the matching results into db. 
A dedicated web page (and a relative API) is built to show the detection rates and some more interesting statistics on your database.

## Running as a Service

If a `./db/picblocksdb.json` exists, you can run

`$ python app.py` 

to spawn a local demo server (`http://127.0.0.1:9001`) to query against.

### Screenshots

Just few screenshots about the initial stage of web user interface 
The submit form. Once you have a given database (`./db/picblocksdb.json`) you can check matching from samples by submitting your samples from 
this form.

<p align="center">
  <img src="static/img/1.png">
</p>

If the submitted sample gets some matches against the given database you should see a block similarity matrix (still under development for a better visualization)

<p align="center">
  <img src="static/img/2.png">
</p>

Finally the matching database statistics generated by the script into `utils/make_stats.py` which takes all the generated block_reports (`block-reports/`) and check them against the generated databases (`./db/picblocksdb.json`) in order to estimate the detection rate on a given database.

<p align="center">
  <img src="static/img/3.png">
</p>

## Contributors

* [Daniel Plohmann](https://github.com/danielplohmann)
* [Marco Ramilli](https://github.com/marcoramilli)
* [Daniele Bellavista](https://github.com/dbellavista)
* [Rony](https://github.com/r0ny123)

## Scores changed in v2.1.0

Up to v2.0.1 the matcher credited a family **once per distinct block hash**, while
`block_bytes` - the denominator every percentage is divided by - counts **every block
occurrence**. A block shared by ten functions therefore contributed its size ten times to
the denominator and once to the numerator, so all four percentages read too low. They now
credit once per matching function, and reported percentages go up accordingly. Measured
over a database built from the Malpedia block reports, matching a sample that is itself in
the database moved from 87.6-99.0% to 98.2-99.4%.

Match reports produced by older versions are not comparable with new ones.

A known residual gap keeps that self-match just under 100%: a report stores the function
ids a block was seen in as a set, so the same hash occurring twice inside one function
counts twice in `block_bytes` but once when scoring. Over the Malpedia block reports this
is ~2.4% of `block_bytes`. Closing it would change the block report format.

## Version History
* 2026-09-13: v2.1.0 - architecture-aware PIC escaping (SMDA >= 4.2.13, AArch64/CIL/Dalvik as well as Intel), corrected matcher scoring (see above), dump/baseaddress routing, report/UI fixes, and a test suite
* 2023-11-24: v2.0.1 - SMDA pinned to 1.12.7 before our bigger fix for PIC calculation
* 2022-09-08: v2.0.0 - (BREAKING CHANGE) now intraprocedural control flow transfers are wildcarded by default, which should improve matching
* 2022-08-04: v1.1.3 - extended format for blockhash representation of functions
* 2021-10-01: v1.1.1 - added script to check detection rates and relative web interface page
* 2021-09-28: v1.1.0 - added simple web user interface and a db connection
* 2021-09-12: v1.0.6 - added submission form fields for bitness and base address to force overrides for those values.
* 2021-08-24: v1.0.5 - improved parsing of bitness from submission filenames.
* 2021-08-20: v1.0.4 - Tweaked result visualization, now showing all unique matches beyond the first 20.
