Skip to content
View fm1320's full-sized avatar

Block or report fm1320

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
fm1320/README.md

Hi, I'm Filip 👋

Machine learning engineer and DevRel working on model inference: making models fast and helping developers use them. Also into philosophy and art, part-time fashion model.

Recent work

At Superlinked, as a founding member of technical staff, here are some public contributions:

  • TopK-Embed-V1 in SIE - I added two multi-vector embedding models (0.8B and 2B) to SIE v0.9.0, Superlinked's open-source inference engine. Fused GPU kernels, CUDA graphs and padding-free batches make them 1.5 to 1.6× faster than the vendor's pipeline. One query takes 6 ms instead of 81 ms. The top search results stay the same. I also fixed a bug that made full batches of large outputs fail in the cluster.
  • SIE - SIE runs small models for agents. I contributed to the project and also led developer growth and adoption and more than doubled its GitHub stars in five months.

I co-authored a paper on FlashNorm.FlashNorm merges the RMSNorm weights into the next linear layer, so transformers run faster.

  • Open source contributions in transformer-tricks, I merged 10+ pull requests. They add GPU benchmarks, Gemma 4 support, and a guide to use FlashNorm with your own model. With FlashNorm, Llama-3.2-1B runs 12.77% faster in Hugging Face Transformers.

Talks and writing

  • Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards - AI Engineer World's Fair 2026 · video
  • The Small Model Infrastructure Nobody Built (So We Did) - AI Engineer Europe 2026 · video
  • One GPU, Four Retrieval Modes: Multi-Model Search Serving - Berlin Buzzwords 2026 · video
  • From BM25 to Mixture-of-Encoders - Haystack EU 2025 · video
  • What Actually Makes Embedding Model Inference Fast? - article

Earlier open source

  • langchain-superlinked - PyPI package with a custom Superlinked mixture-of-encoders retriever for LangChain
  • Superlinked x LlamaIndex - A custom LlamaIndex retriever that uses Superlinked
  • AdalFlow - Contributor to the library for building and optimizing LLM task pipelines: multimodal OpenAI support, the integrations page, and RAG and text-splitter tutorials

Before that

  • Mood-based song recommender - Transformer models and vector search to recommend Spotify songs by mood (Qdrant Vector Space Talks)
  • Data science and ML roles in retail, fintech and biomedical AI

Reach me

LinkedIn, X-Twitter, Substack

Pinned Loading

  1. superlinked/sie superlinked/sie Public

    Open-source inference server and production cluster for all the models your agent needs.

    Python 3.4k 318

  2. OpenMachine-ai/transformer-tricks OpenMachine-ai/transformer-tricks Public

    A collection of tricks and tools to speed up transformer models

    TeX 232 17

  3. song-vibe song-vibe Public

    AI song recommendations based on the feel of a song

    Python 25

  4. superlinked/langchain-superlinked superlinked/langchain-superlinked Public

    Custom superlinked retriever in langchain

    Python 2