Conpack E2E Eval Results

Generated: 2026-07-21 03:39:57 UTC

Summary

No Ctx BEST CONFIDENCE BEST TIME
37.9%
confidence  |  pass rate: 100%
10.5s
avg time  |  p95: 12.2s  |  p99: 12.4s  |  best: 9.2s
MCP Tools WORST
7.9%
confidence  |  pass rate: 0%
12.3s
avg time  |  p95: 18.4s  |  p99: 19.1s  |  best: 7.5s

Comparison

MetricNo CtxMCP Tools
Avg Confidence37.9%7.9%
Pass Rate100%0%
Avg Time10.5s12.3s
Time P9512.2s18.4s
Time P9712.3s18.7s
Time P9912.4s19.1s
Best Time9.2s7.5s
Success Rate3/33/3
Timeouts00
Errors00

Coverage Matrix (Strategy × Vertical)

StrategyNo CtxMCP Tools
fts100% (1/1)0% (0/1)
vector100% (1/1)0% (0/1)
cascade100% (1/1)0% (0/1)

Per-Strategy Recall

StrategyNo CtxMCP Tools
fts100% (1/1)0% (0/1)
vector100% (1/1)0% (0/1)
cascade100% (1/1)0% (0/1)

Per-Query Confidence

QueryStrategyNo CtxMCP Tools
q-fts-1Tokio async runtime schedulerfts35%-
q-vector-1similarity search over high-dimensional embeddingsvector43%-
q-cascade-1Meilisearch typo tolerance rankingcascade35%-
q-fts-2Cache-Control ETag header freshnessfts-14%
q-vector-2in-memory data store with TTL evictionvector-10%
q-cascade-2Result enum pattern matching errorscascade-0%

Query Details

q-fts-1 No Ctx — Tokio async runtime scheduler
Expected: doc-001 — Rust async tokio runtime
Tokio is an async runtime for the Rust programming language, providing the building blocks for writing network applications. It offers a multi-threaded work-stealing scheduler and supports both TCP and UDP protocols.
VerticalDocMatchConfidenceTimeMatched Terms
No Ctxdoc-001YES0.3549.2s
title-partial:3/4tag:asynctag:tokiotag:runtimeterm:runtimeterm:tokio
Agent Responses
No Ctx — recall 100%
To create a Tokio async runtime scheduler for your project, you can follow these specific steps: 1. Install the Tokio async runtime scheduler from the official GitHub repository: ```bash npm install tokio-async @types/async-runtime-scheduler ``` 2. Import the necessary modules in your JavaScript file: ```javascript import { async } from '@types/async-runtime-scheduler'; ``` 3. Implement the scheduler function: ```javascript async function main() { const { getCount, getTasks, setTasks, startScheduling } = await async.runtimeScheduler; const count = getCount(); const tasks = getTasks(); if (count > 0) { setTasks(tasks); } if (tasks.length > 0) { startScheduling(); } } ``` 4. Test the scheduler function using a simple example: ```javascript main(); ``` This code snippet demonstrates how to use the async runtime scheduler to create a simple scheduler that runs tasks on a set schedule.
q-vector-1 No Ctx — similarity search over high-dimensional embeddings
Expected: doc-004 — Vector databases for embeddings
Vector databases like Qdrant, Pinecone, and Milvus store high-dimensional embeddings and enable similarity search using cosine distance, dot product, or Euclidean distance. They power modern RAG and recommendation systems.
VerticalDocMatchConfidenceTimeMatched Terms
No Ctxdoc-004YES0.42910.0s
title-partial:2/3phrase:Euclidean distancetag:vectortag:embeddingsterm:embeddingsterm:euclideanterm:high-dimensionalterm:similarityterm:productterm:vector
Agent Responses
No Ctx — recall 100%
The similarity search over high-dimensional embeddings can be performed using various algorithms, including: 1. K-Nearest Neighbors (KNN): This algorithm finds the most similar documents or points to a given point in the high-dimensional space. It works by finding the k most similar points to the given point and then returns the k most similar documents or points. 2. Cosine Similarity: This algorithm measures the similarity between two vectors based on their dot product and the cosine of their angle. It is a widely used similarity measure and is useful for high-dimensional data. 3. Distance-based Similarity Measures: These measures, such as Euclidean distance, Manhattan distance, and Minkowski distance, are used to compare vectors in high-dimensional spaces. These measures provide a quantitative measure of similarity between vectors, which is useful for high-dimensional data. 4. Mutual Information: This algorithm measures the amount of information that one vector has about another. It is used in various applications, such as feature selection and dimensionality reduction. These algorithms can be used for similarity search over high-dimensional embeddings. For example, a KNN algorithm can be used to find the most similar documents in a high-dimensional space, while a cosine similarity algorithm can be used to compare two vectors and measure their similarity.
q-cascade-1 No Ctx — Meilisearch typo tolerance ranking
Expected: doc-002 — Meilisearch full-text search
Meilisearch is a lightweight, fast, and open-source search engine that delivers instant and relevant search results out of the box. It uses BM25-like ranking and supports typo tolerance, filters, and faceting.
VerticalDocMatchConfidenceTimeMatched Terms
No Ctxdoc-002YES0.35412.4s
title-partial:2/3tag:meilisearchtag:searchterm:toleranceterm:resultsterm:meilisearchterm:ranking
Agent Responses
No Ctx — recall 100%
Meilisearch is a web search engine that provides users with a search interface. Meilisearch typo tolerance ranking is a feature that allows users to adjust the type of characters to which search engines can search. This feature is designed to ensure that users get the most accurate results possible, by using alternative character variations that may be overlooked by conventional search engines. Meilisearch typo tolerance ranking is based on the following technical details: - Meilisearch uses a database of over 200 million web pages to provide users with accurate search results. - Meilisearch uses a search algorithm that is specifically designed for search engines. - Meilisearch uses machine learning algorithms to improve the accuracy of search results. Meilisearch typo tolerance ranking is achieved through the following technical details: - Meilisearch uses a technique called "partial matching" to search for variations of characters in search queries. - Meilisearch uses a technique called "substring matching" to search for variations of characters in search queries. - Meilisearch uses a technique called "fuzzy matching" to search for variations of characters in search queries. Meilisearch typo tolerance ranking is designed to improve the accuracy of search results by reducing the impact of typos and variations in search queries. By using alternative character variations, Meilisearch ensures that users get the most accurate results possible, while also providing a better user experience.
q-fts-2 MCP Tools — Cache-Control ETag header freshness
Expected: doc-005 — HTTP caching headers explained
Cache-Control, ETag, Last-Modified, and Vary headers control how HTTP responses are cached by browsers and CDNs. Cache-Control: max-age=N sets freshness in seconds. Must-revalidate forces rechecking stale entries.
VerticalDocMatchConfidenceTimeMatched Terms
MCP Toolsdoc-005NO0.1437.5s
tag:httptag:cacheterm:cache-controlterm:http
Agent Responses
MCP Tools — recall 0%
To find relevant documents, I will use the conproxy_search command. Here is an example of how to use it: ``` # Usage: conproxy_search "https://www.example.com/cache-control/2015-08-10/ {key} {doc_name} {doc_url}" # Output: ``` In this example, the query parameter `{key}` should be replaced with the desired key. The query parameters `{doc_name}` and `{doc_url}` will be used to retrieve the relevant document(s) from the cache control list. ``` # Output: ``` If the cache control list is not found, the output will be an error message indicating that the key does not exist in the cache control list. If the cache list is found, the output will contain the relevant document(s) with the specified key.
q-vector-2 MCP Tools — in-memory data store with TTL eviction
Expected: doc-003 — Caching strategies in distributed systems
Distributed caches reduce latency and backend load. Strategies include read-through, write-through, write-behind, and cache-aside. TTL and eviction policies determine freshness versus memory pressure.
VerticalDocMatchConfidenceTimeMatched Terms
MCP Toolsdoc-003NO0.09519.3s
tag:performanceterm:memoryterm:systems
Agent Responses
MCP Tools — recall 0%
I can help you search for relevant documents and answer with specific technical details. Question: in-memory data store with TTL eviction Answer: Answer: This is a specific type of data store known as an In-Memory Data Store (IMDS) with a Time-to-Live (TTL) eviction mechanism. IMDS data stores are designed to store data in memory, allowing for fast access, but also requiring periodic flushing to reduce memory usage. In-Memory Data Store (IMDS) with TTL eviction: IMDS data stores are designed to store data in memory, allowing for fast access, but also requiring periodic flushing to reduce memory usage. In-Memory Data Store (IMDS) with TTL eviction is a type of In-Memory Data Store (IMDS) that uses TTL eviction as a mechanism to reduce memory usage by periodically flushing data from the data store. This mechanism is based on the principle of Time-to-Live (TTL), which is used in many databases and data stores to track the expiration time of data. When a data block in a database is updated, the TTL is incremented. After a certain period of time (e.g., one hour), the TTL is decremented, and if the TTL has not expired, the data block is discarded and re-added to the data store. This mechanism helps to reduce memory usage by avoiding unnecessary data storage and freeing up memory for other tasks. In-Memory Data Store (IMDS) with TTL eviction is commonly used in systems that require high performance, low latency, and high availability. This is because it allows for fast access to data by minimizing the amount of memory used and the time it takes to flush data from the data store. It's worth noting that IMDS with TTL eviction is different from traditional databases, which use a data page-by-page approach to data storage. In IMDS with TTL eviction, data is stored in memory as a block, and the TTL is used to determine when the block should be flushed to the disk. This approach provides a trade-off between memory usage and the ability to quickly access data, but it can also lead to increased memory usage over time.
q-cascade-2 MCP Tools — Result enum pattern matching errors
Expected: doc-006 — Rust error handling with Result
Rust uses Result<T, E> for recoverable errors and panic! for unrecoverable ones. The ? operator propagates errors up the call stack. anyhow::Result is a common choice for application code; thiserror for library code.
VerticalDocMatchConfidenceTimeMatched Terms
MCP Toolsdoc-006NO0.00010.0s
Agent Responses
MCP Tools — recall 0%
Error: "conproxy_search.cpp:19:10: error: ‘error’ cannot be found in the target platform. It is likely that the compiler was compiled with a different platform." Error message: ``` error: cannot find ‘error’ in the target platform. It is likely that the compiler was compiled with a different platform. ``` Error description: ``` error: cannot find ‘error’ in the target platform. It is likely that the compiler was compiled with a different platform. ``` The error occurs because the error is from the compiler, not conproxy. It's possible that the compiler was compiled with a different platform, which is why the error is present. Conproxy should be compiled with the target platform's compiler to resolve this issue. To resolve this error, you can try the following steps: 1. Ensure that the compiler's target platform is correctly set in your conproxy.yaml file. 2. Use the conproxy_search tool to search for relevant documents. Here's an example of how you can modify your conproxy.yaml file to set the target platform: ```yaml # Conproxy.yaml # Set the target platform in the conproxy.yaml file # Conproxy.search.yml # Set the target platform in the Conproxy.search.yml file ```