SJ Sai Satyam Jena

Currently ML Engineer at Zynix.ai

Sai Satyam
Jena.

Applied AI Researcher & Machine Learning Engineer

I do AI research and ship it in production. First-author papers at CVPR 2026 and ACL 2026, a national IndiaAI Hackathon win, predictive models serving 1M+ patients, and real-time voice pipelines handling 1,000+ concurrent calls.

Portrait of Sai Satyam Jena
2A* publications
(CVPR · ACL 2026)
1M+patients served by
healthcare ML platform
1,000+concurrent calls,
15+ languages
<700 msend-to-end voice
pipeline latency

Research & recognition

Peer-reviewed publications and national recognition.

  1. [1] CVPR 2026Conference paper DSCA: Dynamic Subspace Concept Alignment for Lifelong Vision-Language Model Editing A lifelong VLM editing framework that updates concepts while preserving previously learned knowledge, addressing catastrophic interference in sequential model edits. Read Paper
  2. [2] ACL 2026Industry Track ResoDiff-44k: High-Fidelity Cross-Lingual Speech and Singing Synthesis via Discrete Diffusion Native 44.1 kHz speech and singing synthesis over codec latents, reducing the cascaded vocoder artifacts common in two-stage pipelines. Read Paper
  3. [3] IndiaAINational winner IndiaAI Hackathon: National Winner Co-developed CricSM AI for critical mineral exploration and integrated NainiV1 to process 60+ years of legacy Geological Survey of India reports. View Announcement
Full publication record See all papers and citations on Google Scholar

Production ML, end to end

From research prototypes to systems that run at scale.

Jun 2025 to PresentBengaluru, India

Zynix.ai

Machine Learning Engineer

  • Developed a heart-disease progression model for a platform serving 1M+ patients, achieving 70%+ held-out accuracy at the 24-month horizon.
  • Architected a production seven-agent Text-to-SQL system using LangGraph.
  • Built multilingual voice AI and neural-memory services supporting 15+ languages and 1,000+ concurrent calls.
  • Designed streaming speech pipelines with barge-in handling and sub-700 ms latency.

May to Nov 2024Bengaluru / Remote

NeuronsAI Pvt. Ltd.

Machine Learning Engineer Intern

  • Designed pipelines for 200+ GB of multimodal mineral-exploration data.
  • Indexed 300+ geological reports, reducing information retrieval from hours to seconds.
  • Integrated LangChain and RAG into the Bhoomi geoscience platform.

Jul to Sep 2023Remote

Exposys Data Labs

Data Engineering Intern

  • Improved price-forecasting accuracy by 20% over baseline models.
  • Built explainable SHAP analyses and Databricks ETL workflows with PySpark and SQL.

Education

IIT (ISM) Dhanbad

Integrated M.Tech in Applied Geophysics · Specialization in AI/ML

2025GPA 8.41/10

Selected projects

Systems built and shipped independently.

Generative video

Distributed Audio-Visual LipSync & Video Dubbing

Multilingual, temporally consistent dubbing without paid GPU infrastructure, using distributed local speech synthesis with Kaggle T4 inference and VRAM-aware preprocessing. Stable within 16 GB VRAM while preserving background and mouth-region consistency.

  • MuseTalk
  • MOSS-TTS
  • FastAPI
  • Gradio
View Demo

Document intelligence

NainiV1

Private analysis of technical PDFs, scans, and tables without third-party APIs. Optimized Qwen2-7B for local reasoning with layout pipelines built on Camelot, img2table, and OpenCV, improving complex table extraction from below 50% to 85%+ within ~14 GB VRAM.

  • Qwen2-7B
  • Camelot
  • img2table
  • OpenCV
View Repository

Agentic systems

AI-Powered PDF Report Generator

A decoupled MCP client-server architecture with asynchronous tools that coordinates research, drafting, and image sourcing, automating the workflow from an initial topic to a structured, illustrated PDF.

  • FastMCP
  • OpenAI API
  • FPDF2
  • Gradio
View Repository

Computer vision

Invoice Data Extraction

An OCR and layout-analysis pipeline that converts scanned invoices into structured XML while retaining layout relationships, reaching 87% key-field precision against an 80% cloud baseline.

  • Python
  • OpenCV
  • OCR
  • XML
View Repository

Technical skills

Depth across the modern ML stack.

Machine Learning & Backend
Python, SQL, PyTorch, Transformers, Scikit-learn, SHAP, FastAPI, AsyncIO, WebSockets
Agentic AI & NLP
LangGraph, LangChain, DSPy, RAG, Text-to-SQL, Qdrant, vector search, prompt orchestration, multi-agent systems
Multimodal AI
STT, TTS, neural audio, lip synchronization, diffusion models, document AI, OCR, video intelligence
Infrastructure
Docker, Kubernetes, ArgoCD, streaming systems, event-driven services, distributed inference
Applied AI
Healthcare ML, clinical risk prediction, value-based care, geospatial ML, explainability

Let's talk.

For research collaborations, roles, or technical discussions.