Scientific quality
Canonical datasets establish numerical parity, non-inferior held-out quality, and coverage of the major representation geometries.
A compact, source-pinned set of quality signals for the first Representax systems paper—covering similarity, pair decisions, frozen probes, clustering, retrieval, multilingual alignment, multimodal retrieval, and JEPA transfer.
This is a research-orientation artifact, not a redistribution of third-party datasets. Dataset adapters should pin source revisions, preserve upstream splits, and stream or cache original artifacts according to their licenses.
Canonical datasets establish numerical parity, non-inferior held-out quality, and coverage of the major representation geometries.
Paired jobs on identical hardware establish throughput, utilization, capacity, compilation, and time-to-quality claims.
The goal is to measure the main signals, not reproduce every MTEB task. Each geometry receives a named evaluator and an explicit primary metric.
| Signal | Evaluator | Initial datasets | Primary metric |
|---|---|---|---|
| Graded symmetric similarity | SimilarityEvaluator | STSBenchmark.v2, SICK-R | Cosine Spearman |
| Pair decisions | PairClassificationEvaluator | Sprint Duplicate Questions | Maximum AP across declared similarities |
| Frozen linear separability | ClassificationProbeEvaluator | Banking77.v2 | Accuracy; macro F1 secondary |
| Unsupervised global geometry | ClusteringEvaluator | Twenty Newsgroups.v2 | V-measure |
| Text retrieval | InformationRetrievalEvaluator | SciFact, NFCorpus, ArguAna | nDCG@10; MRR/MAP/recall secondary |
| Multilingual retrieval | InformationRetrievalEvaluator | MIRACL | Macro-language and per-language nDCG@10 |
| Cross-lingual alignment | Ordinary IR/qrels adapter | FLORES-200 | F1 or retrieval accuracy |
| Image ↔ text | InformationRetrievalEvaluator | Flickr30k | Recall@1/5/10 and nDCG@10 |
| Audio ↔ text | InformationRetrievalEvaluator | AudioCaps | Hit-rate@5 and recall@1/5/10 |
| Video ↔ text | InformationRetrievalEvaluator | MSR-VTT | Recall@1/5/10 and nDCG@10 |
| JEPA transfer | JEPARepresentationEvaluator + frozen probe | ImageNet-1K | Linear-probe top-1; k-NN/collapse secondary |
SICK-R complements STS-B with controlled compositional examples. MIRACL and FLORES make multilingual behavior a first-class requirement. FLORES uses the ordinary query/corpus/qrels geometry—there is no special “translation evaluator.”
The encoder is frozen. Representax embeds train, validation, and test examples once; fits only a lightweight linear classifier on train embeddings; selects probe hyperparameters on validation data; and reports held-out labels. This measures how linearly accessible task information is, not how well the encoder can be fine-tuned.
Pin normalization, classifier family, regularization grid, class weighting, solver tolerance, maximum iterations, seed, and selection split. Identical embeddings must reproduce the maintained reference evaluator.
Short excerpts show each data geometry and label meaning. The linked, revision-pinned upstream artifact remains authoritative; these excerpts must not become packaged fixtures.
5,749 train · 1,485 validation · 1,362 test after duplicate removal.
9,927 test pairs.
101,000 test pairs; 1,000 positives.
9,993 train · 3,076 test.
Ten deterministic 1,000-document clustering samples.
1,109 queries · 5,183 abstracts · 339 test qrels.
3,237 queries · 3,633 documents · 12,334 test qrels.
1,406 queries · 8,674 candidate arguments.
997 development · 1,012 devtest aligned records.
1,000 images · 5,000 captions · 5,000 qrels.
883 audio queries · 4,411 captions · 4,411 qrels.
879 test videos with audio and captions.
Freeze the backbone, extract the declared representation, fit a pinned linear probe on train, and report validation top-1.
Licenses marked non-commercial or unspecified require review before examples become durable public fixtures. The evaluator tests should use synthetic/checkpointed oracle values unless redistribution is clearly permitted.
STS-B remains relevant in 2026 for historical continuity and as a fast test of graded symmetric similarity. Sentence-BERT and SimCSE reported it directly; current model papers increasingly emphasize MTEB/MMTEB category and aggregate scores. It cannot stand in for retrieval, clustering, multilingual, or multimodal quality.
Sentence Transformers does not define universal systems numbers. It supplies maintained implementations, evaluator semantics, and useful recipes. Representax must publish the exact matched jobs and rerun both systems on identical hardware, software, examples, ordering, shapes, precision, optimizer, schedule, batch, and update count.
Published model-card scores can validate quality plumbing. Published throughput from different hardware cannot establish a systems claim.