Integrations
Selected systems, pipelines, and libraries highlighting ML infrastructure and engineering depth. Focus is entirely on model serving latency, quantization performance, vector retrieval, and operational stability.
-
ApplyRail
Job portal matching 50K+ postings to international students using Gemini embeddings + pgvector; serves from a single Proxmox node at ~180ms p95.
Flask Postgres pgvector Gemini Docker180msp95 Latency -
KernServe
Custom LLM inference runtime optimized for 6GB GPUs utilizing custom CUDA-accelerated kernels, dynamic batching, and INT4 quantization to achieve 45+ tokens/sec on local consumer hardware.
Python CUDA C++ GGUF PyTorch vLLM45+tokens/sec -
RAG-Eval
Enterprise RAG observability platform calculating semantic similarity, retrieval precision, and answer faithfulness metrics in real-time for Mastercard Assistant Platform (MAP), handling 1.2M monthly queries.
Python Databricks Splunk Datadog SageMaker1.2Mqueries/mo -
Bio-Segment
Edge-deployed insect mortality detection model running on GCP IoT Core and SageMaker pipelines, automating computer vision analysis of 100K+ daily farm images with an 18% harvest yield gain.
AWS GCP Step Functions Airflow BigQuery Dataflow18%yield gain