SLM Embeddings Server
Starts a local CPU-optimized embedding server to compute dense document and query vectors on standard hardware.
🚀 Overview & Capabilities
Starts a local CPU-optimized embedding server to compute dense document and query vectors on standard hardware.
Key Features
- Loads quantized mini-LM or BGE embeddings locally
- High-speed cosine similarity index built directly in memory
- Provides local HTTP API endpoint for integration
- Under 200 MB RAM memory usage footprint during idle states
💻 Installation
Install the local CPU-optimized package using pip:
# Install local CPU-optimized package
pip install slm-embeddings
🐙 Checkout from GitHub
Clone only this agent's folder from the monorepo using Git sparse-checkout — no need to download the full repository:
Option 1 — Sparse Checkout (Recommended)
Option 2 — Full Repository Clone
💡 Tip: After checkout, install the package locally with pip install -e ./slm_embeddings to run in editable mode without publishing to PyPI.
⚙️ Configuration API
Constructor Parameters
Instantiate SLMEmbeddingsServer with performance options:
| Parameter | Type / Default | Description |
|---|---|---|
| model_path | str | None | Explicit path to ONNX model weights. If omitted, downloads standard checkpoints. |
| cache_dir | str | None | Directory to store model weights offline. Defaults to ~/.cache/slm-embeddings/. Also settable via SLM_EMBEDDINGS_SERVER_CACHE_DIR. |
| n_threads | int | 4 | CPU thread count for ONNX inference. Optimize for CPU core count. Also settable via SLM_EMBEDDINGS_SERVER_N_THREADS. |
Methods
| Method Signature | Return Type | Description |
|---|---|---|
embed(texts, system_prompt=None, user_input=None) | list[float] | Generates dense vectors from standard string lists. |
Method Parameters (Execution Customization)
All main execution methods accept optional system routing parameters:
| Parameter | Type / Default | Description |
|---|---|---|
| system_prompt | str | None | Optional custom system prompt instruction to override the default system template response parameters. |
| user_input | str | None | Optional additional user-supplied target text variables or contextual keys. |
Quick Start
from slm_embeddings import SLMEmbeddingsServer
server = SLMEmbeddingsServer()
vector = server.embed(
["sample test"],
system_prompt="Calculate semantic weights",
user_input="Normalized cosine distance"
)
print(vector)
Environment Variables
Configure agent parameters globally using environment values:
| Environment Variable | Default | Purpose |
|---|---|---|
| SLM_EMBEDDINGS_SERVER_N_THREADS | 4 | Sets CPU inference execution threads. |
| SLM_EMBEDDINGS_SERVER_CACHE_DIR | ~/.cache/slm-embeddings/ | Default directory to store downloaded ONNX weights. |
CPU Performance Tuning
To run the SLMEmbeddingsServer engine efficiently on CPU under 1.5 GB memory footprint:
- Match Threads to Core Count: Set
n_threadsorSLM_EMBEDDINGS_SERVER_N_THREADSto match the physical CPU core count. - Sequential Processing: Avoid concurrent processing when batch files are large.
- Garbage Collection: Clear variables and run
gc.collect()to release model RAM blocks after execution.
Verified Input & Output Logs
Diagnostic execution console response running locally on CPU:
→ INPUT:
"sample test"
← OUTPUT:
"Vector dimension check: 1024"