Skip to content
View amareshhebbar's full-sized avatar

Organizations

@onenot8

Block or report amareshhebbar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
amareshhebbar/README.md

Amaresh Hebbar

Founding AI Engineer @ Impossible AI · Applied AI Engineer @ AnacodicAI Labs
Agentic LLM Systems · Model Auditing · Post Training & Alignment · Multi Agent Infrastructure · Medical AI
I design, fine tune, audit, and ship production grade AI systems end to end.

LinkedIn Email Impossible AI AnacodicAI Labs Hugging Face PyPI W&B LeetCode ORCID


About

I build agentic AI systems: multi agent pipelines, LLM infrastructure, model auditing and interpretability tooling, on device inference, and domain fine tuning, and I take them all the way to production.

  • Founding AI Engineer at Impossible AI. Building the on device LLM inference layer with cloud fallback routing for an agentic AI fitness platform, keeping 80% of sessions fully offline at 45ms baseline latency.
  • Applied AI Engineer at AnacodicAI Labs (founded at Boston University). Self volunteered on ClinicalSearch, a multi agent clinical evidence retrieval system for plastic and reconstructive surgery.
  • Open source author. Shipped Modeldiffr to PyPI: audits exactly what changed between a base LLM and any fine tune, quant, merge, or edit, with paired deltas and 95% bootstrap confidence intervals. Also shipped TrueNorth (LLM infrastructure engine, 1,258 tests, 8 provider routing, about 90% cost reduction), BitNarrow (zero training weight surgery on 4 bit LLMs), and GitGrounded (AI regression testing).
  • Upstream contributor. Fixed a critical Transformers 5.5+ import crash in unslothai/unsloth zoo (PR #897).
  • Fine tuning at scale. Published a 16 model medical AI suite (Qwen2.5) covering ICD 10, CPT, DRG coding, SNOMED mapping, clinical NLP, PM JAY classification, and Hindi medical. Every model is trained with a QLoRA → DoRA → ORPO → merge pipeline on a real, paired SFT dataset (also published).
  • Led and mentored a 10 person engineering team across frontend, backend, and mobile, delivering two concurrent AI product lines.
  • Hackathons. SANS FIND EVIL! (DFIR Automation) · Google Cloud Rapid Agent (GitLab Partner) · INDIA RUNS (Redrob AI × Hack2Skill, Data & AI) · BITSoM Vertex Builders Pitch Fest (Top 150 of 2,300, Software Automation AI Track).
  • Research grade rigor. Published benchmarks (100% precision on SANS DFIR triage), open SFT datasets, statistically grounded model diffs, and W&B tracked training runs.
  • Based in Bengaluru, India · Open to remote first AI engineering roles (IST, comfortable with US and EU overlap).

Experience

Role Organisation What I do
Founding AI Engineer
Mar 2025 to Present
Impossible AI · Website
Bengaluru, India (Remote)
Building the on device LLM inference layer with cloud fallback routing for the Impossible AI agentic fitness platform, keeping 80% of sessions fully offline at 45ms baseline latency. Orchestrating 16 specialised fine tuned models behind a unified domain classification layer, running the Supabase backend at zero errors under peak load, and mentoring 10 engineering interns
Applied AI Engineer (Research & Development)
Aug 2026 to Present
AnacodicAI Labs (founded at Boston University, nonprofit) · GitHub · Agentic Cookbook
Self volunteered, contributing under project lead Rashan Kaur
ClinicalSearch: multi agent clinical evidence retrieval for plastic and reconstructive surgery. 4 specialised agents (Search, Medical Fact Checker, Synthesizer, Evaluator) on the AWS Strands Agents SDK, Pinecone hybrid search over 2.5M+ clinical abstracts, a FastAPI backend, and a Vite and TypeScript frontend. Built a fully local Ollama inference stack that cut query latency by 38%, a critique loop that lifted Top K recall by 22%, and benchmark suites for extraction quality and LLM provider comparison

Featured on LinkedIn: view post


Tech Stack


Featured Projects

Project What it does Stack Highlights
Modeldiffr
PyPI
Audits what changed between a base LLM and any fine tune, quant, merge, or edit. One command runs both models under identical conditions and reports what moved, by how much, and whether it beats evaluation noise Python · PyTorch · Transformers Flagship · paired deltas with 95% bootstrap CIs · token level KL divergence · CI gate mode · roadmap to crosscoder and causal diffing · BitNarrow as layer 2
BitNarrow
PyPI · Weights
Zero training, in place weight surgery on 4 bit quantized LLMs using Winsorized activation profiling and Gram Schmidt orthogonal projection across the residual stream Python · PyTorch · bitsandbytes · Unsloth Full spectrum projection across o, gate, up, down layers · 95th percentile outlier clamping · 386MB hot swap weight patch · 0% refusal with 100% logic retention · Abliteration Weights collection
TrueNorth
PyPI
Developer first LLM infrastructure engine: declare the outcome in YAML, it owns the full multi turn conversation lifecycle through a 13 stage pipeline Python · TS · Go · RN 1,258 tests · 4 SDKs on PyPI + NPM · hallucination firewall (94%) · 8 provider routing · about 90% cost reduction
GitGrounded
PyPI
Catches AI regressions before users do: diffs a prompt or model change, has an AI write targeted tests, judges old vs new, returns PASS, WARN, or FAIL Python · Claude · Groq · Ollama · Streamlit Git mode and live endpoint mode · version history per API · automatic PR comments · 22 tests in CI · BITSoM Vertex Builders Pitch Fest (Top 150 of 2,300)
Medical AI Suite
Datasets
16 fine tuned Qwen2.5 specialist models for medical coding, billing and clinical NLP QLoRA · DoRA · ORPO · Unsloth · HF 13 published models + 16 open SFT datasets · <1% token hallucination · >99% structural format compliance · live demo · Apache 2.0
LogPoseSIFT
Devpost
Autonomous DFIR orchestrator: MCP server wraps 200+ SANS SIFT tools as typed Go endpoints Go · Claude · Gemini · MCP · Volatility 3 100% precision · 92.8% recall · 0 hallucinations · SANS FIND EVIL! Hackathon · extended into AllBlue for Splunk
ShiftLeft
Devpost
Autonomous 5 agent bug fixing pipeline: reads repo → triages → generates fix → opens MR Python · LangGraph · Gemini · GitLab MCP End to end in about 60s, zero human steps · Google Cloud Rapid Agent Hackathon
layerFourth (private repo) Fully local autonomous AI web agent: headed Chromium via raw CDP (no Playwright or Selenium), AI driven mouse and keyboard control, 4 layer extraction fallback (DOM → Accessibility Tree → Network sniff → Vision OCR) Python · CDP · Vector DB 134/134 tests passing across 12 build phases · dual layer vector memory (ephemeral + persistent)
More projects
Project What it does Stack Highlights
HireSignal
Live sandbox
Ranks 100K candidates against a Senior AI Engineer JD in about 35s on CPU: multi signal scoring, honeypot detection, semantic embeddings Python · sentence transformers · NumPy No GPU, no API, no network during ranking · 85 honeypots caught · 10 tests · INDIA RUNS Hackathon
PocketLLM 100% offline Android AI chat running LLMs on device via a MediaPipe C++ bridge React Native · Expo · MediaPipe C++ · AWS S3 9 open weight models (0.4 to 5.2 GB) · prompts never leave the device
raiseTicket / IssueLoop AI managed ticket queue for open source repos: test run failures become LLM triaged tickets, fix proposals, re tests, and escalations Python · Supabase · Ollama Local first embeddings and reasoning · pluggable provider config
OceanAI Website Investor grade 3D product site for the OceanAI health platform Next.js · TypeScript 65 files, 28 routes · live Claude API playground demos embedded
HatPet Custom Linux desktop pet: transparent, borderless, always on top window that wanders the screen in 8 directions, idles, and responds to drag Godot 4 · GDScript XWayland compatible movement (native Wayland blocks window repositioning) · built from scratch, no pet framework

Fine tuned models live on Hugging Face · Packages on PyPI (modeldiffr · truenorth-framework · bitnarrow · gitgrounded) · Training runs tracked on Weights & Biases


Open Source Contributions

Repository Contribution Impact
unslothai/unsloth zoo PR #897: resolved a critical ModuleNotFoundError import crash Restored runtime stability for LLM training workflows on Transformers 5.5+
anacodicAI labs ClinicalSearch multi agent retrieval, local Ollama inference stack, and the Agentic Cookbook Nonprofit clinical AI research founded at Boston University

Hugging Face: Medical AI Fine Tuned Model Suite

A suite of Qwen2.5 specialist models, one per clinical task. Each model is trained through a consistent QLoRA → DoRA → ORPO → merge pipeline (via Unsloth + TRL) on a dedicated, published SFT dataset, with no synthetic training data. Released under Apache 2.0; training tracked on W&B.

Collection: Medical AI Fine Tuned Model Suite · Datasets: AxisMapper Medical AI Suite · Code: AxisMapper

Model Size Task Dataset (rows) Method GPU
icd10-coder-qwen25-7b 7B Clinical text → ICD 10 CM code + justification icd10-coder-sft (74.7k) QLoRA → DoRA → ORPO → merge A40 48GB
icd10-coder-qwen25-7b-merged 8B Merged full weights build of the ICD 10 coder (no adapter load) icd10-coder-sft (74.7k) QLoRA → DoRA → ORPO → merge A40 48GB
snomed-mapper-qwen25-7b 7B Clinical concept → SNOMED CT mapping snomed-mapper-sft (74.7k) QLoRA → DoRA → ORPO → merge A40 48GB
clinical-summarizer-qwen25-7b 7B Clinical note summarization clinical-summarizer-sft (30k) QLoRA → DoRA → ORPO → merge A40 48GB
medical-billing-qwen25-3b 3B Medical billing code generation medical-billing-sft (17k) QLoRA → DoRA → ORPO → merge A40 48GB
cpt-coder-qwen25-3b 3B Procedure text → CPT code cpt-coder-sft (17k) QLoRA → DoRA → ORPO → merge A40 48GB
radiology-coder-qwen25-3b 3B Radiology report → diagnostic code radiology-coder-sft (25.1k) QLoRA → DoRA → ORPO → merge A40 48GB
pmjay-classifier-qwen25-3b 3B India PM JAY scheme package classification pmjay-classifier-sft (11.1k) QLoRA → DoRA → ORPO → merge A40 48GB
discharge-qa-qwen25-3b 3B QA over discharge summaries discharge-qa-sft (30k) QLoRA → DoRA → ORPO → merge A40 48GB
medical-ner-qwen25-3b 3B Clinical named entity recognition medical-ner-sft (16.7k) QLoRA → DoRA → ORPO → merge A40 48GB
hindi-medical-qwen25-3b 3B Hindi language medical assistant hindi-medical-sft (19.7k) QLoRA → DoRA → ORPO → merge A40 48GB
icd10-to-drg-qwen25-1b 1.5B ICD 10 → DRG for reimbursement grouping icd10-to-drg-sft (5.39k) QLoRA → DoRA → ORPO → merge A40 48GB
insurance-classifier-qwen25-1b 1.5B CPT/HCPCS → Stark Law DHS classification insurance-classifier-sft (1.6k) QLoRA → DoRA → ORPO → merge A40 48GB
ayurveda-icd-qwen25-1b 1.5B Ayurveda term → ICD mapping ayurveda-icd-sft (3k) QLoRA → DoRA → ORPO → merge A40 48GB
pharmacy-ner-qwen25-1b 1.5B Pharmacy and drug entity recognition pharmacy-ner-sft (3.5k) QLoRA → DoRA → ORPO → merge A40 48GB

Pipeline (all models): Qwen2.5 Instruct base → QLoRA SFT (4 bit NF4, rank 16, α 32) → DoRA → ORPO preference alignment → adapter merge. Optimizer paged_adamw_8bit · cosine schedule, LR 0.0002 · trained on RunPod (NVIDIA A40 48GB) via Unsloth + TRL. Per model wall time ranges from about 0.4 h (1.5B configs) to about 1.9 h (7B configs). Training tracked at wandb.ai/amareshhebbar. Datasets built from authoritative real world sources (for example CMS FY2026 ICD 10 CM and the HCPCS Stark Law DHS list), not LLM generated.

Live demos (Spaces): icd10 coder demo · hiresignal


Hackathons

Submission Hackathon Track What it does
GitGrounded · PyPI BITSoM Vertex × H2S Builders Pitch Fest 2026 (Top 150 of 2,300) Software Automation AI Tests an AI app before and after a prompt or model change, AI written targeted tests, AI judge, PASS, WARN, or FAIL verdict
ShiftLeft · Repo Google Cloud Rapid Agent GitLab Partner Label a GitLab issue → autonomous 5 agent pipeline reads the repo, triages the bug, writes the fix, and opens an MR in under 60 seconds
Poneglyphs: ShiftLeft Google Cloud Rapid Agent GitLab Partner Label a GitLab issue shiftleft → 5 agent pipeline reads GitLab Orbit, triages, writes fix, opens MR
LogPoseSIFT · Repo SANS FIND EVIL! DFIR Automation Autonomous DFIR orchestrator: deploys an AI crew via strict MCP endpoints, runs SIFT diagnostics, triages and self corrects in seconds
AllBlue SANS FIND EVIL! DFIR Automation Splunk alerts trigger autonomous AI forensic triage, with IOC findings pushed back as structured events. 100% precision, 0 hallucinations
HireSignal · Live sandbox INDIA RUNS · Redrob AI × Hack2Skill Data & AI Challenge Ranks 100K candidates against a Senior AI Engineer JD in about 35s on CPU: multi signal scoring, 85 honeypots detected, per candidate reasoning

GitHub Stats

GitHub Stats

  </td>
  <td>

Top Languages

  </td>
</tr>

GitHub streak


Activity

This month's commit summary

Monthly commit and pull request activity

Today vs yesterday commit and pull request activity

Auto generated from live account data, refreshed every 6 hours · Last updated: 2026-10-06 12:22 UTC


Highlights


Open to remote first AI engineering roles: LLM infrastructure, agentic systems, model evaluation, fine tuning, or AI product engineering.
[email protected] · LinkedIn
Let's build something intelligent.

Profile views

Pinned Loading

  1. TrueNorth TrueNorth Public

    Declare the outcome, skip the logic. The developer-first infrastructure engine for structured, multi-turn AI conversations with built-in safety, async follow-ups, and native compliance

    Python 24

  2. ShiftLeft ShiftLeft Public

    Autonomous GitLab bug-fixing agent — reads repo, triages issues with Gemini 3.1 Pro, writes fix, opens MR. Built with LangGraph + GitLab MCP.

    Python

  3. LogPoseSIFT LogPoseSIFT Public

    An autonomous, multi-agent DFIR orchestrator. LogPose utilizes custom MCP boundaries to safely execute SIFT tools and synthesize breach data into actionable timelines at machine speed

    Go 1

  4. AllBlue AllBlue Public

    Autonomous multi-agent DFIR orchestrator — Splunk alerts trigger AI triage, findings pushed back to Splunk. 100% precision, 0 hallucinations. Claude + SIFT + Go MCP Server.

    Go

  5. hiresignal hiresignal Public

    Ranks 100K candidates against a Senior AI Engineer JD in ~35s on CPU — multi-signal scoring, honeypot detection, semantic embeddings. INDIA RUNS Hackathon Track 1.

    Python