SLM Data Analyst
Loads local CSV, Parquet, or Excel files. Answers statistical questions, performs calculations, and auto-generates data visualization code.
🚀 Overview & Capabilities
Loads local CSV, Parquet, or Excel files. Answers statistical questions, performs calculations, and auto-generates data visualization code.
Key Features
- Direct pandas dataframe parsing and stats calculator
- Translates user query into python matplotlib/pandas code blocks
- Generates summary tables and column distribution charts
- 100% offline analysis of highly sensitive company sheets
💻 Installation
Install the local CPU-optimized package using pip:
# Install local CPU-optimized package
pip install slm-data
🐙 Checkout from GitHub
Clone only this agent's folder from the monorepo using Git sparse-checkout — no need to download the full repository:
Option 1 — Sparse Checkout (Recommended)
Option 2 — Full Repository Clone
💡 Tip: After checkout, install the package locally with pip install -e ./slm_data to run in editable mode without publishing to PyPI.
⚙️ Configuration API
Constructor Parameters
Instantiate SLMDataAnalyst with performance options:
| Parameter | Type / Default | Description |
|---|---|---|
| model_path | str | None | Explicit path to ONNX model weights. If omitted, downloads standard checkpoints. |
| cache_dir | str | None | Directory to store model weights offline. Defaults to ~/.cache/slm-data/. Also settable via SLM_DATA_ANALYST_CACHE_DIR. |
| n_threads | int | 4 | CPU thread count for ONNX inference. Optimize for CPU core count. Also settable via SLM_DATA_ANALYST_N_THREADS. |
Methods
| Method Signature | Return Type | Description |
|---|---|---|
analyze_file(csv_path, query, system_prompt=None, user_input=None) | dict | Parses tables, checks datatypes, and generates mathematical statistical logs. |
Method Parameters (Execution Customization)
All main execution methods accept optional system routing parameters:
| Parameter | Type / Default | Description |
|---|---|---|
| system_prompt | str | None | Optional custom system prompt instruction to override the default system template response parameters. |
| user_input | str | None | Optional additional user-supplied target text variables or contextual keys. |
Quick Start
from slm_data import SLMDataAnalyst
analyst = SLMDataAnalyst()
result = analyst.analyze_file(
"sales.csv",
"summarize sales",
system_prompt="Prioritize revenue aggregations",
user_input="Limit charts to bar plots"
)
print(result)
Environment Variables
Configure agent parameters globally using environment values:
| Environment Variable | Default | Purpose |
|---|---|---|
| SLM_DATA_ANALYST_N_THREADS | 4 | Sets CPU inference execution threads. |
| SLM_DATA_ANALYST_CACHE_DIR | ~/.cache/slm-data/ | Default directory to store downloaded ONNX weights. |
CPU Performance Tuning
To run the SLMDataAnalyst engine efficiently on CPU under 1.5 GB memory footprint:
- Match Threads to Core Count: Set
n_threadsorSLM_DATA_ANALYST_N_THREADSto match the physical CPU core count. - Sequential Processing: Avoid concurrent processing when batch files are large.
- Garbage Collection: Clear variables and run
gc.collect()to release model RAM blocks after execution.
Verified Input & Output Logs
Diagnostic execution console response running locally on CPU:
→ INPUT (CSV):
{"file": "sales.csv", "query": "summarize sales"}
← OUTPUT:
{
'columns': [],
'summary': 'Calculated total revenue by region: East ($15,000), West ($22,000).'
}