QuickClassify Benchmark

Frozen pretrained embeddings + task-specific logistic heads. Evaluated 2026-09-29 08:47:59.

Architecture: BAAI/bge-large-en-v1.5 (1024d frozen encoder, ~1.2 GB) → StandardScaler → LogisticRegression (C=0.01, newton-cg)
Each head: 1 scaler + 1 linear classifier. No fine-tuning, no MLP, no generation.
Training regime: Each head trained on up to 10,000 examples from the dataset's training split (seed=42 for subsampling).
Comparison context: QuickClassify heads are task-trained; Jeff and similar zero-shot models are not. This is a deployment comparison, not an equal-supervision experiment.
Baselines: TF-IDF+LR tuned via 3-fold CV over C in {0.01, 0.1, 1, 10}, analyzer in {word, char_wb}, ngram_range in {(1,1), (1,2)}, max_features in {10000, 30000}. Vectorizer fit inside each fold.

Results

TaskClassesTest NSplit AccuracyMacro F1 MajorityTF-IDF+LRJeff Latency (p50)Head Size
sms_spam
SMS spam detection
2 1115 test 99.1%
[0.986, 0.996]
98.0% 86.6% 98.7% not evaluated 109ms
emb 108 + clf 0.22
29 KB
dbpedia
Wikipedia article ontology classification
14 2000 test 96.0%
[0.951, 0.968]
95.9% 7.6% 96.5% not evaluated 114ms
emb 114 + clf 0.22
81 KB
imdb
Movie review sentiment (positive/negative)
2 2000 test 94.8%
[0.939, 0.957]
94.8% 52.0% 88.3% not evaluated 252ms
emb 252 + clf 0.22
29 KB
banking77
Banking customer service intent detection
77 2000 test 94.3%
[0.933, 0.953]
94.3% 1.4% 90.0% 19.5% 118ms
emb 114 + clf 2.25
334 KB
ag_news
News article topic classification
4 2000 test 90.5%
[0.892, 0.918]
90.5% 25.2% 89.0% 89.5% 79ms
emb 79 + clf 0.22
41 KB
sst2
Movie review sentiment (positive/negative)
2 872 validation 90.1%
[0.882, 0.921]
90.1% 50.9% 79.5% 84.5% 104ms
emb 104 + clf 0.22
29 KB
clinc_oos
Intent detection with out-of-scope
151 2000 test 88.4%
[0.871, 0.898]
92.1% 17.9% 78.6% not evaluated 118ms
emb 116 + clf 2.08
632 KB
massive_intent
Amazon MASSIVE voice command intents
60 2000 test 88.1%
[0.867, 0.895]
86.4% 7.0% 83.0% not evaluated 114ms
emb 113 + clf 0.40
271 KB
tweet_eval_offensive
Offensive language detection
2 860 test 81.0%
[0.785, 0.836]
74.8% 72.1% 79.0% not evaluated 105ms
emb 105 + clf 0.23
29 KB
tweet_eval_emotion
Tweet emotion detection
4 1421 test 78.1%
[0.760, 0.804]
74.7% 39.3% 66.7% not evaluated 105ms
emb 105 + clf 0.22
41 KB
emotion
Text emotion detection
6 2000 test 75.5%
[0.736, 0.774]
67.8% 34.8% 87.2% not evaluated 102ms
emb 102 + clf 0.22
49 KB
tweet_eval_sentiment
Tweet sentiment analysis
3 2000 test 66.2%
[0.640, 0.683]
65.7% 47.5% 54.2% not evaluated 89ms
emb 88 + clf 0.23
37 KB
snli
Natural language inference
3 2000 test 65.6%
[0.634, 0.675]
65.2% 33.7% 48.9% not evaluated 102ms
emb 101 + clf 0.22
37 KB
Average (13 tasks) 85.2%83.9%
Jeff comparison: Not yet evaluated locally. Jeff is a zero-shot decision model (no per-task training). A fair comparison requires running Jeff on identical evaluation examples with the same label meanings. QuickClassify's advantage comes from task-specific training; Jeff's advantage is generalization without training data.

Task-Specific Notes

SST-2:
Evaluated on validation split; official test labels are not public.
SMS Spam:
Random train/test split (test_size=0.2, seed=42); no standard benchmark split.
SNLI:
Input encoded as "premise [SEP] hypothesis". Label -1 (unlabeled) filtered.
CLINC-OOS:
151 classes including out-of-scope. In-scope accuracy 96.5%, OOS detection 51.7%.
MASSIVE:
English subset only (config='en'). String labels used directly.

Infrastructure

Hardware:
macOS-15.6.1-arm64-arm-64bit, 16 cores, 68.7 GB RAM
Encoder:
BAAI/bge-large-en-v1.5, ~1.2 GB on disk, 2.15s load time
Total pack:
1666 KB (13 heads, no encoder)
Threads:
Default (sentence-transformers manages threading)
Batch size:
1 (single-example inference measured)
Caching:
No embedding cache; each request embeds fresh
Truncation:
Default sentence-transformers truncation (512 tokens)

Reproduce

# Install
pip install sentence-transformers scikit-learn fastapi uvicorn

# Build model pack from source datasets
python -m quickclassify.serve.build_pack --out data/model_pack

# Run evaluation
python -m quickclassify.serve.evaluate --baselines --latency

# Start server with playground
python -m quickclassify.serve.server
# Open http://localhost:8400

API

# Predict with a pretrained head
curl -X POST http://localhost:8400/v1/predict \
  -H "Content-Type: application/json" \
  -d '{"text": "I was charged twice", "task": "banking77"}'

# List capabilities
curl http://localhost:8400/v1/capabilities

# Jeff-compatible endpoint (matches criteria to pretrained heads)
curl -X POST http://localhost:8400/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"model":"quickclassify","state":"I was charged twice",
       "questions":{"intent":{"type":"choice","criteria":{
         "transaction_charged_twice":null,"request_refund":null,
         "cancel_transfer":null}}}}'