Consulting

AI engineering help for teams that need it to work in production.

I am a Principal AI Engineer who builds LLM systems end to end: retrieval and search, fine-tuning on private data, and the multi-node training and serving infrastructure underneath. I take on a small number of consulting engagements where that experience shortens the path from prototype to something you can rely on.

The write-ups on this site are the portfolio. If the problems in them look like yours, we should talk.

You might be here because

  • Your RAG prototype impressed the demo audience and disappointed the first real users, and nobody can say precisely why.
  • Prompting has hit its ceiling. You suspect fine-tuning would help, but the training run keeps dying and the GPU bill keeps growing.
  • Your team is strong but has never shipped an LLM feature. You need senior judgement on architecture and vendors, not another headcount.

How I can help

Four shapes of engagement. Each starts with a free call and a written proposal with a fixed price. If your problem does not fit any of these, describe it anyway.

LLM architecture review

Fixed scope · about two weeks

A written assessment of the system you have or the one you are planning: retrieval design, evaluation, cost and latency, failure modes, and the order to fix them in. You leave with a prioritized plan your team can execute with or without me.

Fit: Best when a prototype works in a demo and nobody is sure it will hold up with real users.

Retrieval and search systems

Scoped project · typically four to eight weeks

Hybrid search (BM25 plus dense vectors plus reranking), agentic RAG that routes queries to specialized retrievers, and the evaluation harness that tells you whether a change helped. Built on your stack, handed over with documentation.

Fit: Best when answers are inconsistent, retrieval quality is guesswork, or a specialized domain such as healthcare needs more than vanilla RAG.

Fine-tuning and distributed training

Scoped project

Supervised fine-tuning and DPO on open-weight models (Qwen, LLaMA, DeepSeek and others), multi-node training with DeepSpeed, torchrun or Ray, plus the checkpointing, NCCL networking and observability that keep a long run alive.

Fit: Best when prompting has hit its ceiling, inference costs need to come down, or your data cannot leave your infrastructure.

Fractional AI lead

Monthly retainer

A standing weekly working session plus asynchronous review of designs, pull requests and vendor decisions. Senior judgement for a team that is building AI features without a senior AI engineer on staff.

Fit: Best for a small engineering team that needs direction more than headcount.

Selected work

Each of these links to a technical write-up on this site. The training runs include the parts that went wrong. I would rather you judge the work than a testimonial.

DeepSeek-V3 multi-node training in BF16

Multi-node BF16 training runs of the 671B-parameter DeepSeek-V3 mixture-of-experts model across three and four A100 nodes: DeepSpeed and Ray Train orchestration, NCCL networking, driver and fabric-manager stability, checkpoint hygiene and Prometheus/Grafana monitoring. Written up as an eight-part series plus a technical overview of the model.

DeepSpeedRayNCCLObservability

Qwen 2.5-72B two-stage fine-tuning

Supervised fine-tuning with ZeRO-3 CPU offload and LoRA, followed by Direct Preference Optimization on preference data, for a content-generation workload.

SFTDPOLoRA

Agentic RAG and hybrid search architecture for medical search

How I approach query routing to specialized agents over a lexical-plus-semantic retrieval stack with neural reranking and diversity optimization, plus the data pipelines that feed it. The architecture I would start from for a domain-specific search product, written up in full.

RAGHybrid searchHealthcare

Marketing mix modeling platform

Architecture and lessons from building Bayesian MMM as a product rather than a PDF: asynchronous GPU training, budget optimization, and a UI that turns posteriors into recommendations non-technical stakeholders can act on.

BayesianPlatformMarketing

LLaMA 3.1 8B distributed fine-tuning

Sixteen-GPU fine-tuning with DeepSpeed ZeRO-3 offload and LoRA, with stable loss curves and efficient memory use.

DistributedDeepSpeed

Image model fine-tuning (SDXL, Flux)

Fine-tuning and serving diffusion models for image generation workloads, including the technical details most write-ups skip.

DiffusionSDXLFlux

How an engagement works

  1. 1

    A 30-minute call

    You describe the problem; I ask the questions that decide whether I can help. No charge, no pitch deck.

  2. 2

    A written proposal

    Scope, deliverables, timeline and a fixed price, in writing, before any work starts. If a smaller engagement would get you there, I say so.

  3. 3

    Delivery in the open

    Working code in your repositories, a weekly demo, and documentation your team can run with after I leave.

About me

I am Thomas Kalnik, a Principal AI Engineer and programming generalist. I have built AI systems in healthcare, digital marketing and the creator economy. Before that I worked in data engineering and data science at Criteo, 1upHealth, Dynamo Software and Wellington Management. That background matters: most production AI problems turn out to be data and infrastructure problems wearing a model costume.

I hold a Master's in Computer Science from the University of Illinois Urbana-Champaign, an MBA and MS in Finance from Northeastern University, and a Bachelor's from SUNY Geneseo. The finance training shows up in how I scope work: I care about the cost of a system and the business outcome it is supposed to move, not just whether the model is clever.

More on the about page, and in the technical posts.

Questions I usually get

What kind of companies do you work with?
Mostly product companies with an engineering team that is shipping, or trying to ship, an LLM-based feature: healthcare, marketing technology and the creator economy are the industries I know best, but the engineering problems travel well.
Do you work remotely?
Yes. All engagements are remote, with regular video working sessions and asynchronous review in between. On-site visits can be arranged for kickoff or a workshop if that helps your team.
Will you work inside our existing codebase and infrastructure?
That is the default. Work happens in your repositories and on your cloud accounts, so nothing is locked in a consultant-owned environment.
Who owns the work?
You do. All code, models and documentation produced during an engagement belong to you. I am happy to sign your NDA before the first call.
How many clients do you take on?
A small number at a time, so each engagement gets real attention. If I am not available when you reach out, I will tell you when I will be, or point you to someone who fits.
What do engagements cost?
Fixed-scope projects and monthly retainers; every one starts with a free call and a written, fixed-price proposal. The budget field on the form helps me suggest the right shape of engagement.

Start a conversation

Tell me what you are building and where it is stuck. I reply personally within two business days, usually with a few questions and a suggested next step. Prefer email? kalnik.thomas@gmail.com.

No newsletter, no follow-up sequence. Just a reply from me.