Frozen pretrained embeddings + task-specific logistic heads. Evaluated 2026-09-29 08:47:59.
| Task | Classes | Test N | Split | Accuracy | Macro F1 | Majority | TF-IDF+LR | Jeff | Latency (p50) | Head Size |
|---|---|---|---|---|---|---|---|---|---|---|
| sms_spam SMS spam detection |
2 | 1115 | test | 99.1% [0.986, 0.996] |
98.0% | 86.6% | 98.7% | not evaluated | 109ms emb 108 + clf 0.22 |
29 KB |
| dbpedia Wikipedia article ontology classification |
14 | 2000 | test | 96.0% [0.951, 0.968] |
95.9% | 7.6% | 96.5% | not evaluated | 114ms emb 114 + clf 0.22 |
81 KB |
| imdb Movie review sentiment (positive/negative) |
2 | 2000 | test | 94.8% [0.939, 0.957] |
94.8% | 52.0% | 88.3% | not evaluated | 252ms emb 252 + clf 0.22 |
29 KB |
| banking77 Banking customer service intent detection |
77 | 2000 | test | 94.3% [0.933, 0.953] |
94.3% | 1.4% | 90.0% | 19.5% | 118ms emb 114 + clf 2.25 |
334 KB |
| ag_news News article topic classification |
4 | 2000 | test | 90.5% [0.892, 0.918] |
90.5% | 25.2% | 89.0% | 89.5% | 79ms emb 79 + clf 0.22 |
41 KB |
| sst2 Movie review sentiment (positive/negative) |
2 | 872 | validation | 90.1% [0.882, 0.921] |
90.1% | 50.9% | 79.5% | 84.5% | 104ms emb 104 + clf 0.22 |
29 KB |
| clinc_oos Intent detection with out-of-scope |
151 | 2000 | test | 88.4% [0.871, 0.898] |
92.1% | 17.9% | 78.6% | not evaluated | 118ms emb 116 + clf 2.08 |
632 KB |
| massive_intent Amazon MASSIVE voice command intents |
60 | 2000 | test | 88.1% [0.867, 0.895] |
86.4% | 7.0% | 83.0% | not evaluated | 114ms emb 113 + clf 0.40 |
271 KB |
| tweet_eval_offensive Offensive language detection |
2 | 860 | test | 81.0% [0.785, 0.836] |
74.8% | 72.1% | 79.0% | not evaluated | 105ms emb 105 + clf 0.23 |
29 KB |
| tweet_eval_emotion Tweet emotion detection |
4 | 1421 | test | 78.1% [0.760, 0.804] |
74.7% | 39.3% | 66.7% | not evaluated | 105ms emb 105 + clf 0.22 |
41 KB |
| emotion Text emotion detection |
6 | 2000 | test | 75.5% [0.736, 0.774] |
67.8% | 34.8% | 87.2% | not evaluated | 102ms emb 102 + clf 0.22 |
49 KB |
| tweet_eval_sentiment Tweet sentiment analysis |
3 | 2000 | test | 66.2% [0.640, 0.683] |
65.7% | 47.5% | 54.2% | not evaluated | 89ms emb 88 + clf 0.23 |
37 KB |
| snli Natural language inference |
3 | 2000 | test | 65.6% [0.634, 0.675] |
65.2% | 33.7% | 48.9% | not evaluated | 102ms emb 101 + clf 0.22 |
37 KB |
| Average (13 tasks) | 85.2% | 83.9% |
# Install
pip install sentence-transformers scikit-learn fastapi uvicorn
# Build model pack from source datasets
python -m quickclassify.serve.build_pack --out data/model_pack
# Run evaluation
python -m quickclassify.serve.evaluate --baselines --latency
# Start server with playground
python -m quickclassify.serve.server
# Open http://localhost:8400
# Predict with a pretrained head
curl -X POST http://localhost:8400/v1/predict \
-H "Content-Type: application/json" \
-d '{"text": "I was charged twice", "task": "banking77"}'
# List capabilities
curl http://localhost:8400/v1/capabilities
# Jeff-compatible endpoint (matches criteria to pretrained heads)
curl -X POST http://localhost:8400/v1/systemone \
-H "Content-Type: application/json" \
-d '{"model":"quickclassify","state":"I was charged twice",
"questions":{"intent":{"type":"choice","criteria":{
"transaction_charged_twice":null,"request_refund":null,
"cancel_transfer":null}}}}'