Skip to main content

Open-Source Ecosystem Roadmap

Engineering roadmap for our open-source developer tooling: LLM Context Forge and Velox GTM.


1. LLM Context Forge (Context & Token Infrastructure)​

Deterministic, zero-telemetry context budgeting, chunking, and cost estimation across LLM providers.

✅ Completed (v0.2.x)​

  • Exact Multi-Backend Tokenization: Pure deterministic encoding with cl100k_base, o200k_base, and HuggingFace models.
  • Intelligent Document Chunking: 5 chunking strategies (Sentence, Paragraph, Semantic/Heuristic, Code, Fixed) with strict parameter validation.
  • Priority Context Packing: ContextWindow algorithm mimicking system-level queues with overflow dropping.
  • Multimodal Vision Token Sizing: VisionTokenCounter calculating exact pixel-tile token costs for OpenAI (GPT-4o detail tiers), Anthropic (Claude 3.5 1568px bounds), and Google Gemini (768px patches).
  • Live Dynamic Pricing Index: Integrated APIPriceIndexProvider consuming API Price Index (600+ models with CC BY 4.0 attribution) and Cheapest LLM API fallback with 1-hour disk TTL caching (~/.cache/llm_context_forge/).
  • CLI & API Server: llm-context-forge count, chunk, assemble, pricing lookup, and FastAPI /docs service.

🚧 In Progress (v0.3 – v0.4)​

  • Streaming Context Window: Async generator packing (stream_context_window()) enabling dynamic client-side token delivery.
  • KV-Cache State Tracking: Tracking prompt prefixes against Anthropic Prompt Caching and OpenAI prefix caching to predict cache-hit vs. cache-miss costs.
  • Framework Integrations: Pinned optional dependency extras for LangChain (ContextForgeTextSplitter) and LlamaIndex (ContextForgeNodeParser).

📋 Planned (v1.0 – v2.0)​

  • Rust FFI Core (forge-core): Native Rust core via PyO3 (Python) and napi-rs (TypeScript) for >2,000,000 tokens/sec chunking and packing.
  • Multi-Modal Native Audio & Video: Native token estimations for audio spectrograms and video frames.
  • WASM Edge Compilation: Edge-ready builds (wasm32-unknown-unknown) for Cloudflare Workers and Vercel Edge.

2. Velox GTM (Sales-to-Ops Automation Engine)​

Lightweight GTM orchestration, offline lead scoring, and automated sales-to-ops handoff.

✅ Completed (v2.1.0)​

  • Architecture Modernization: Modular src/velox_gtm/ layout with clean packaging configuration in pyproject.toml.
  • Zero-Setup Offline ICP Engine: 3-axis lead scoring matrix (LinkedInGTMLeadEngine) running with sub-millisecond local latency and velox --demo.
  • Automated Notion ERP Provisioning: Dual relational database deployment (Pipeline Tracker ↔ Client Vault & Project Delivery).
  • Google Docs Proposal Compiler: Markdown AST parser and formatter compiling blueprints into styled Google Docs.
  • Comprehensive Mock Test Suite & CI: 100% offline pytest test harness with Notion and Google API mocks; GitHub Actions multi-version CI across Python 3.10, 3.11, and 3.12.

🚧 In Progress (v2.2.0)​

  • FastAPI Webhook Listener: Background daemon (velox listen --port 8000) to directly ingest incoming webhooks from n8n, Stripe, and HubSpot.
  • Pluggable Local CRM Adapters: SQLite and PostgreSQL backends for sovereign, self-hosted operational tracking without Notion dependencies.

📋 Planned (v2.5.0 – v3.0.0)​

  • Dynamic Scoring Rule Engine: Configurable YAML/JSON rule schema (icp_rules.yaml) replacing hardcoded scoring parameters.
  • Distributed RevOps Processing: Async worker queue (Redis / Celery) for high-throughput batch CSV lead qualification.
  • Audit & PII Compliance: Data sanitation engine and audit logging for SOC2/GDPR compliance across CRM handoffs.