S·R

AI / ML engineering · Coimbatore, India

Sai Reddy Annem

I build AI systems that keep working when the data doesn't cooperate — missing sensors, imbalanced labels, noisy live feeds — and I ship them with tests, a one-command launcher, and an honest list of what's still broken.

Third-year B.Tech, Computer Science & Engineering (AI), Amrita Vishwa Vidyapeetham · graduating 2028. Currently building VoxAgent, a multilingual AI phone receptionist for Indian small businesses.

Checking build status on GitHub…

sai@portfolio: ~

6public repos, CI green on every one
318automated tests across them
2IEEE papers (1 published, 1 in review)
40/40simulated drone landings, 2.9 cm mean error
§ 01

Selected work

Six projects, each with a figure you can play with. Every number on this page comes from the repository it links to.

01

Multi-hazard disaster risk, explained

Bayesian networks · pgmpy · FastAPI · live weather feeds

An 18-node discrete Bayesian network turns one district's weather and river readings into flood, landslide, cyclone and heatwave probabilities — with a plain-language explanation, a confidence statement, and an action chosen by maximum expected utility. It runs on live Open-Meteo and GloFAS feeds with no API key.

The interesting result isn't the headline AUC. With every sensor reporting, gradient boosting and logistic regression win. Take sensors away and the BN degrades gracefully, because it marginalises what it didn't see instead of imputing it.

  • 35,072 real district-days used to re-learn the weather layer
  • 85 tests · team of four, course 22AIE301
Held-out ROC-AUC. Drag the slider to knock out sensor readings.
02

InterviewAce AI

FastAPI · React + TypeScript · SQLAlchemy · speech-to-text · LLMs

Upload a resume, get the score an applicant-tracking system would give it, sit a voice mock interview generated from your own projects, and get a week-by-week study plan built from the gaps that interview exposed.

Design rule: every AI feature has a real offline fallback. Add an Anthropic, OpenAI or Gemini key and questions are written from your resume; leave it empty and a curated bank plus a heuristic scorer take over — and the header badge always says which one is running.

  • 92 tests, no network or API key needed
  • ATS score is arithmetic: fixed weights, every point traceable
InterviewAce resume analyzer showing a 98/100 ATS score

Where the 100 points come from — hover a bar

Skill match · 40 — do you have the skills the role asks for

Real screens from the running app, using the fictional sample resume in the test fixtures.
03

Vision-guided landing on moving pads

MuJoCo · OpenCV · AprilTags · PID + feed-forward · quaternions

A simulated quadrotor flies to a numbered address, finds the right pad with its own downward camera, tracks it while it moves, and lands. Frames are really rendered and real tag36h11 markers are really decoded — when the drone tilts, the tag genuinely swings across the image.

The base paper's fixed-gain PID never lands on a moving pad. Step through the controller changes to see what actually fixed it.

  • 40/40 Monte Carlo landings · 1.2 cm static, 6.8 cm moving
  • 20/20 mid-flight diverts, 0 go-arounds

Ablation on a moving pad · 6 seeds each

    Simulation only: idealised motors, no latency model beyond the camera rate. Team course project, built on our team's original MuJoCo model.
    04

    Interpretable eye-disease diagnosis

    PyTorch · prototype learning · CBAM attention · ODIR-5K

    One forward pass over a fundus photo gives eight disease probabilities, a localisation map per disease, and a short text report naming the region the evidence came from. One branch is designed and trained from scratch with a CIPL-style prototype head; a pretrained ResNet-18 + CBAM branch measures what that constraint costs.

    Biggest lesson: on a heavily imbalanced dataset, loss weighting mattered more than architecture. Plain BCE collapses to "no disease" for everything.

    Fundus image, attention map and per-disease localisation
    Validation-split numbers with an image-level split, so likely optimistic. The repo's harness now does patient-level train/val/test; re-running the grid is next.
    05

    GreenScale Cloud

    Flask · scikit-learn · ECDSA · Merkle trees · scheduling

    A carbon-aware multi-cloud scheduler: a gradient-boosted model predicts how a job will behave, then the scheduler picks the region and start hour that minimise a weighted mix of carbon, energy, cost and latency. Every decision is written as a signed transaction into a Merkle-summarised, proof-of-work ledger — edit one number and validation points at the exact record.

    My part: the scheduler core and the experiment harness. Pick a policy and watch the trade-off: the pure-carbon policy wins the table and loses the customer.

    400-job simulated trace, seed 42. Baseline is "home region, start now". Ledger overhead: 2.7 ms per transaction.
    06

    polylag

    asyncio · httpx · websockets · risk management

    An event-driven system built to measure whether prediction markets lag the news: it watches feeds, matches headlines against hand-written rules, checks whether a market has not yet repriced, and only then acts — in watch or paper mode by default.

    The README opens with "this will probably lose money", and the tooling is built to tell you when the edge is dead. Click a layer to see what it guards.

    • 68 tests · no profitability claimed anywhere

    § 02

    Now building: VoxAgent

    A multilingual AI receptionist for Indian clinics, salons and coaching institutes — answers calls in English, Telugu and Hindi, books appointments, logs every conversation. Private while it's a product; built in stages so each one works before the next starts.

    1. Stage 1Text receptionist with real auto-booking into Google Sheets
    2. Stage 2Voice: speech-to-text → the same agent → neural TTS
    3. Stage 2.5Live, interruptible calls with server-side voice activity detection
    4. Stage 3Real phone numbers via Twilio, business portal, eval suite

    Architecture choices I'd defend in an interview: one llm.py seam so the model provider is swappable; static facts in the prompt and live data only through tool calls — no RAG, because a salon's price list doesn't need retrieval.

    § 03

    Path

    From lab assignments to systems with tests.

    1. Started B.Tech CSE (AI)

      C, Python, signal processing, maths for computing.

    2. First real ML

      Job-scheduling simulator, AI demand response, and a Q-learning agent for 4×4 Tic-Tac-Toe that became an IEEE ESCI 2026 paper.

    3. Research & data work

      Hybrid CNN–ViT image forgery detection (IEEE Access, in review); IoT respiratory monitoring; DBMS knowledge portal.

    4. Started VoxAgent

      From a terminal chatbot to real phone calls in under three weeks of commits.

    5. Systems with tests

      Disaster-risk BN, GreenScale Cloud, InterviewAce, polylag, the drone landing simulation and the eye-disease work — each with CI on GitHub.

    § 04

    Stack

    Only what I've shipped something with. Hover a tool to see where it was used.

    Learning Agent tooling (MCP servers, LangGraph) · prototype-based interpretability · evaluation methodology for small, imbalanced medical datasets.

    § 05

    Research

    § 06

    Open to AI/ML and backend internships,
    and to building something useful together.