Portfolio

Gokkulnath T S · ML Systems Engineer · Qualcomm AI Accelerator

LLM serving, tuned for the hardware under the hood.

inference-session · vllm
$ 

About

Machine Learning Systems Engineer, solving inference at scale.

I'm a Machine Learning Systems Engineer at Qualcomm, working on the Qualcomm AI Accelerator product line. My job is the unglamorous part of LLM inference: taking serving stacks like vLLM, LLM-D, and LMCache — frameworks written around GPU assumptions — and making them behave on specialized silicon. That means reworking how they lay out memory, schedule batches, and place the KV cache on an accelerator that was designed for the problem but not for their code. The constraints are the interesting part: adapt, don't assume, and keep the whole thing stable at production scale.

At Ericsson I built an NLP log-analytics tool that surfaced insights from telecom log streams, and automated testing for SIP-based Toll Free and IVR services. Between the two I spent a semester as a research assistant at IISc, working on adversarial perturbations with conditional GANs.

My research roots are in signal processing and machine learning: a 2016 Springer-published paper on extended Kalman filtering for missile-interception tracking, published work on efficient SVM classification, and a streak of adversarial-ML and computer-vision experiments that still shape how I think about failure modes.

Outside of work I play table tennis, fix broken electronics, watch movies — and I'm almost always programming.

Experience

Where I've built things.

ML Systems Engineer — Qualcomm

Jul 2021 — Present
  • Serving high-performance LLM inference on the Qualcomm AI Accelerator.
  • Adapting vLLM / LLM-D / LMCache-based serving stacks to specialized AI hardware.

Research Assistant — IISc

Jan 2021 — Jun 2021
  • Explored adversarial perturbations using Conditional GANs — targeted, class-aware perturbation generation as an extension of NAG.

Software Development Engineer — Ericsson

Sep 2017 — Jan 2020
  • Automated testing and tooling for SIP-based Toll Free and IVR services.
  • Built VNF lifecycle test automation for cloud-native telecom workloads.

Software Development Engineer Intern — Ericsson

Apr 2017 — Sep 2017
  • Developed an NLP log-analytics tool to surface insights from telecom log streams.

Selected Work

LLM-era engineering, in production.

ML Systems · Qualcomm

LLM inference serving on Qualcomm AI Accelerator

Serving LLMs on the Qualcomm AI Accelerator: adapting open-source serving stacks to specialized accelerators, and shipping inference that holds up at production scale.

ML Systems · Qualcomm

LMCache on the Qualcomm AI Accelerator

Adapting and extending LMCache, an open-source KV-cache acceleration layer, to the Qualcomm AI Accelerator serving stack — cutting redundant KV recomputation across requests so long-context serving stays fast and memory-efficient.

ML Systems · Qualcomm

vLLM & LLM-D on specialized hardware

Adapting high-throughput serving frameworks — vLLM and LLM-D — to the Qualcomm AI Accelerator.

Earlier work in adversarial ML and computer vision: NAG (adversarial examples for deep nets), Face Aging with CycleGAN, and a clustered-SVM approach for large-scale classification.

Earlier explorations

From my computer-vision days.

A few projects from before I moved to LLM serving, with walkthroughs on my YouTube channel. My current focus is inference engineering — see the work section for that.

Archive · Computer Vision

Mosquito Detection: YOLO Object Detection on Mobile

A mobile-first YOLO object detector that spots mosquitoes on camera in real time, built to see how far lightweight detection can go on constrained edge hardware.

Watch on YouTube
Archive · Adversarial ML

NAG: Network Adversary Generator

A PyTorch implementation that generates adversarial examples to probe the robustness of deep neural networks, exploring how small imperceptible perturbations can mislead modern classifiers.

Watch on YouTube

Publications

Selected papers.

A Novel Clustered Support Vector Machine with Reduced Support Vectors for Big Data Classification

Gokkulnath T.S., Ramanathan R · TechRxiv preprint

Reducing the number of support vectors cuts model complexity, enabling real-time applications on low-power computing devices with lighter hardware requirements.

Read preprint

Earlier work: "Tracking Inbound Enemy Missile for Interception from Target Aircraft Using Extended Kalman Filter" — SSCC 2016, Springer. Full list on Google Scholar →

© 2026 Gokkulnath T S · Built with Jekyll & Chirpy