Open-Source Ecosystem Roadmap
Engineering roadmap for our open-source developer tooling: LLM Context Forge and Velox GTM.
1. LLM Context Forge (Context & Token Infrastructure)
Deterministic, zero-telemetry context budgeting, chunking, and cost estimation across LLM providers.
✅ Completed (v0.2.x)
- Exact Multi-Backend Tokenization: Pure deterministic encoding with
cl100k_base,o200k_base, and HuggingFace models. - Intelligent Document Chunking: 5 chunking strategies (Sentence, Paragraph, Semantic/Heuristic, Code, Fixed) with strict parameter validation.
- Priority Context Packing:
ContextWindowalgorithm mimicking system-level queues with overflow dropping. - Multimodal Vision Token Sizing:
VisionTokenCountercalculating exact pixel-tile token costs for OpenAI (GPT-4o detail tiers), Anthropic (Claude 3.5 1568px bounds), and Google Gemini (768px patches). - Live Dynamic Pricing Index: Integrated
APIPriceIndexProviderconsuming API Price Index (600+ models with CC BY 4.0 attribution) and Cheapest LLM API fallback with 1-hour disk TTL caching (~/.cache/llm_context_forge/). - CLI & API Server:
llm-context-forge count,chunk,assemble,pricing lookup, and FastAPI/docsservice.
🚧 In Progress (v0.3 – v0.4)
- Streaming Context Window: Async generator packing (
stream_context_window()) enabling dynamic client-side token delivery. - KV-Cache State Tracking: Tracking prompt prefixes against Anthropic Prompt Caching and OpenAI prefix caching to predict cache-hit vs. cache-miss costs.
- Framework Integrations: Pinned optional dependency extras for LangChain (
ContextForgeTextSplitter) and LlamaIndex (ContextForgeNodeParser).
📋 Planned (v1.0 – v2.0)
- Rust FFI Core (
forge-core): Native Rust core via PyO3 (Python) and napi-rs (TypeScript) for >2,000,000 tokens/sec chunking and packing. - Multi-Modal Native Audio & Video: Native token estimations for audio spectrograms and video frames.
- WASM Edge Compilation: Edge-ready builds (
wasm32-unknown-unknown) for Cloudflare Workers and Vercel Edge.
2. Velox GTM (Sales-to-Ops Automation Engine)
Lightweight GTM orchestration, offline lead scoring, and automated sales-to-ops handoff.
✅ Completed (v2.1.0)
- Architecture Modernization: Modular
src/velox_gtm/layout with clean packaging configuration inpyproject.toml. - Zero-Setup Offline ICP Engine: 3-axis lead scoring matrix (
LinkedInGTMLeadEngine) running with sub-millisecond local latency andvelox --demo. - Automated Notion ERP Provisioning: Dual relational database deployment (
Pipeline Tracker↔Client Vault & Project Delivery). - Google Docs Proposal Compiler: Markdown AST parser and formatter compiling blueprints into styled Google Docs.
- Comprehensive Mock Test Suite & CI: 100% offline
pytesttest harness with Notion and Google API mocks; GitHub Actions multi-version CI across Python 3.10, 3.11, and 3.12.
🚧 In Progress (v2.2.0)
- FastAPI Webhook Listener: Background daemon (
velox listen --port 8000) to directly ingest incoming webhooks from n8n, Stripe, and HubSpot. - Pluggable Local CRM Adapters: SQLite and PostgreSQL backends for sovereign, self-hosted operational tracking without Notion dependencies.
📋 Planned (v2.5.0 – v3.0.0)
- Dynamic Scoring Rule Engine: Configurable YAML/JSON rule schema (
icp_rules.yaml) replacing hardcoded scoring parameters. - Distributed RevOps Processing: Async worker queue (Redis / Celery) for high-throughput batch CSV lead qualification.
- Audit & PII Compliance: Data sanitation engine and audit logging for SOC2/GDPR compliance across CRM handoffs.