All services
Technology capabilities

LLM Fine Tuning

Baaz fine-tunes open-weight models so a small model you own performs like a frontier model on your task, at a fraction of the cost per token. We work in parameter-efficient tuning, LoRA and 4-bit QLoRA on open weights from 7B to 70B, engineered for deterministic output your code can parse every time. Because the weights are self-hosted, training and serving can run in your own data centre or fully air-gapped, which is often the only route when sensitive data cannot reach a hosted API. And because we modernise the systems these models feed, the tuned model lands in a working application instead of a notebook.

Fine-tuning is worth doing when a model keeps getting the shape of the answer wrong: the wrong format, the wrong vocabulary, the wrong level of caution, the wrong reasoning style. Prompting and retrieval fix a lot of that, and we will tell you when they are enough. Past that point the fastest route is to train the behaviour into the weights. We come at this as a software delivery problem. The measure of success is a model running inside your product under load, returning output your code can parse every time.

Our offerings in LLM fine tuning

Each of these ships with an evaluation set, a held-out regression suite, and a documented recipe you can rerun without us.

Supervised Fine-Tuning

Train on curated input and output pairs to lock in response format, domain vocabulary, tone, and structured output that parses reliably on the first attempt.

LoRA & QLoRA Adapters

Parameter-efficient training that updates a small set of weights instead of the whole model. For larger open weights we run 4-bit QLoRA with NF4 quantisation, double quantisation, and paged optimizers, which is what brings 7B to 70B models within reach of sensible hardware.

Deterministic Output Engineering

Training sets built against a strict JSON or SQL schema, with deliberate abstention examples so the model returns a clean null for an attribute it cannot find instead of inventing a plausible one. This is usually what separates a demo from something you can put behind an API.

Preference & Reward Optimisation

DPO, ORPO, KTO, and GRPO for the behaviours supervised examples cannot express: which of two acceptable answers is better, when to refuse, when to ask instead of guessing.

Distillation Into Small Models

Move the behaviour of a large model into a model in the 1B to 8B range that you own and serve. For high-volume structured work this is where the cost curve changes by an order of magnitude.

Training Data Curation

Mining real production failures, deduplicating, decontaminating against your test set, and generating synthetic coverage for gaps. A few thousand well-scored examples beat tens of thousands of scraped ones.

Evaluation & Regression Gates

A golden dataset built from real failures, deterministic checks first and a calibrated judge second, wired into CI so a regression blocks the release instead of reaching users.

Integration Into Live Systems

We build and modernise ERPs, CRMs, logistics engines, and computer-vision stacks, so the tuned model goes into an actual workflow with the error handling, retries, fallbacks, and audit trail that implies.

Technology stack & training infrastructure

The frameworks, methods, and serving targets we use to take an open-weight model from a base checkpoint to a monitored production endpoint.

Open-weight base models

Families we tune and benchmark against each other, from 7B task models up to 70B.

  • Llama
  • Mistral
  • Qwen
  • Gemma
  • Phi

Training frameworks

Single-GPU speed for iteration, distributed training when the run outgrows one card.

  • Axolotl
  • Unsloth
  • TRL
  • LLaMA-Factory

Parameter-efficient tuning

The 4-bit path that makes large open weights trainable without a cluster.

  • LoRA
  • QLoRA
  • NF4 + double quantisation
  • Paged optimizers
  • PEFT

Alignment methods

Preference and reward optimisation for the behaviours examples cannot state.

  • DPO
  • ORPO
  • KTO
  • GRPO

Distributed compute

Sharded training across multiple GPUs, with long-context runs where the task needs them.

  • PyTorch FSDP
  • DeepSpeed
  • H100 / A100 clusters
  • Sequence parallelism

Serving & inference

Throughput-oriented serving with per-request adapter selection and quantised weights.

  • vLLM
  • TensorRT-LLM
  • Multi-LoRA serving
  • AWQ / GPTQ / FP8

Output validation

The layer that keeps a malformed response from reaching your database.

  • JSON Schema
  • Constrained decoding
  • SQL contract tests
  • Abstention checks

Evaluation & observability

The measurement layer that decides whether a checkpoint ships.

  • Golden datasets
  • LLM-as-judge (calibrated)
  • CI regression gates
  • Drift monitoring

When fine-tuning is the right call

Fine-tuning is the wrong tool about as often as it is the right one. These are the signals we look for before recommending it, and if none of them hold we will say so.

The problem is form, not facts

Fine-tuning reliably changes how a model answers. It does not reliably teach it new facts. If your gap is knowledge that changes weekly, retrieval is the fix and we will build that instead.

You need output your code can trust

A response that is right nine times in ten is a bug when the tenth goes into a database. Schema-strict tuning with abstention is the reliable route to parseable output and an honest null.

You already have failure traces

The best training set is the log of what your current model got wrong in production. Teams with that history get a useful model in weeks. Teams without it need instrumentation first.

Volume makes cost real

At high request volume on a narrow task, a tuned small model changes the unit economics and the tail latency. At low volume, a frontier API is cheaper than the GPUs and the engineering.

A pilot has stalled short of production

Plenty of AI work dies between a promising notebook and a shipped feature. If that is where yours is, the missing piece is usually engineering, not more modelling.

The data cannot leave your servers

Some data simply cannot be sent to a hosted API, and no contract makes that acceptable: patient records under HIPAA, card and transaction data under PCI, personal data with residency conditions attached, defence and critical-infrastructure work, or the drawings, formulas, and source code that are the business itself. A hosted frontier model means every prompt crosses your boundary to a third party who logs it, and an internal security review will stop that. A tuned open-weight model is the way out, because the weights run on hardware you control, inside your own network or fully air-gapped, and the sensitive text never leaves the building.

The model has to stay put

Hosted models get versioned, deprecated, and quietly changed underneath you, which is a problem when a regulator asks you to reproduce a decision from eighteen months ago. A model you hold as a file does not move unless you move it.

How a fine-tuning engagement runs

Evals Before Training

We agree the pass mark first and build the harness to measure it. Training without an eval you trust produces a model nobody can approve.

Failure Mining

We work through your production traces to separate prompting problems, retrieval problems, and genuine behaviour problems. Only the third kind needs a fine-tune.

Schema Definition

We pin down the exact JSON or SQL contract the model has to satisfy, including how it should signal that a value is genuinely absent.

Dataset Construction

Curate, score, and deduplicate real examples, write the abstention cases explicitly, generate synthetic coverage for the gaps, and hold back a slice that training never sees.

Base Model Selection

We benchmark open-weight candidates from 7B to 70B against your eval, because which base wins changes from task to task and only the eval can tell you.

Adapter Training

LoRA or 4-bit QLoRA runs with tracked hyperparameters, checkpoints, and cost per run, so any result can be reproduced or rolled back later.

Preference Tuning

Where judgement matters, a preference pass over ranked outputs teaches the model which acceptable answer is the better one.

Regression Gate

The candidate runs against the held-out suite plus general-capability slices, so we catch a model that learned your task and forgot how to reason.

Wiring Into the Product

The model goes behind a service in your actual application, with schema validation, retries, fallbacks, and logging, so a bad response degrades instead of corrupting data.

Serving & Drift Watch

Deploy behind a versioned endpoint with adapter hot-swapping, then keep scoring live traffic so quality drift surfaces before your users report it.

Why choose Baaz for LLM fine tuning?

We have built software since 2018 from Bengaluru with a US office in Sheridan, Wyoming, and we treat fine-tuning as production engineering. The training run is the short part. Everything around it decides whether the model is safe to ship.

Evaluation first, always

The harness exists before the first training job. You get a number you can defend internally, not a demo that looks better than it measures.

We will talk you out of it

If prompting, retrieval, or a better base model solves your problem, that is the recommendation. Fine-tuning you did not need is the most expensive kind.

Deterministic by construction

Strict schemas and explicit abstention examples are part of the dataset from day one, so a missing attribute comes back as a null your code can handle instead of a confident guess.

It ships into the system you run

We modernise ERPs, CRMs, logistics engines, and vision stacks for a living, so the model lands inside a real workflow with the plumbing that requires.

Your data stays inside your boundary

Training runs in your cloud account or ours under your terms, with open weights, no lock-in to a hosted tuning service, and an air-gapped path when you need one.

Handover, not dependency

You get the weights, the dataset, the recipe, and the eval harness. Retraining next quarter should not require another statement of work.

How an engagement is shaped

Most work starts small on purpose. A pilot answers whether tuning helps your task at all before anyone commits to a pipeline, and the eval harness it produces is what the larger engagement is judged against.

  • Scoped validation pilot

    One task, one base model, one adapter. We build the eval harness, assemble a first dataset from your traces, train, and report the measured gap against your current baseline. The deliverable is a decision backed by numbers, plus everything needed to go further.

  • End-to-end production pipeline

    Dataset pipeline, schema and abstention design, adapter and preference training, regression gates in CI, serving with hot-swappable adapters, integration into the application that consumes it, and drift monitoring after launch. Scoped per task, per model size, and by how much of your existing system has to change.

Engagements in between are common and scoped the same way. If your problem turns out not to need fine-tuning, the pilot is where we tell you.

Where the tuned model runs

A fine-tuned open-weight model is an artifact you control, so your constraints decide where it runs. We serve it wherever the data is allowed to be, and keep the base model shared so several adapters can run behind one endpoint.

  • Your own VPC on AWS, GCP, or Azure
  • On-premise GPU hardware
  • Air-gapped and offline environments
  • Managed endpoint operated by Baaz
  • Multi-adapter serving on a shared base model
  • Versioned rollback to any prior checkpoint

LLM Fine Tuning - Frequently Asked Questions

LLM fine tuning continues training an existing language model on your own examples so its behaviour changes: the format it answers in, the vocabulary it uses, the reasoning style it follows, and when it declines. The result is a model you own that is specialised to your task, rather than a general model you steer with a long prompt on every request.

Usually both, in that order. Retrieval is the right tool when the gap is knowledge, especially knowledge that changes, because you can update a document store instantly and fine-tuning does not reliably teach new facts. Fine-tuning is the right tool when the gap is behaviour that a prompt keeps failing to enforce. The sequence we recommend is prompt, then retrieval, then fine-tune, then distil, stopping as soon as the eval is satisfied.

By training for abstention on purpose. The dataset is built against a strict JSON or SQL schema and includes examples where an attribute is genuinely absent and the correct answer is a null, so returning nothing becomes a learned behaviour instead of a failure mode. On top of that, schema validation and constrained decoding sit in front of your application, so a malformed response is caught before it reaches your data.

With a scoped validation pilot: one task, one base model, one adapter. We build the evaluation harness first, assemble a dataset from your own traces, train, and report the measured gap against your current baseline. That gives you a decision backed by numbers before anyone commits to a full pipeline, and the harness carries forward into the larger engagement if you go ahead.

Less than most teams expect, if the examples are good. A few hundred to a few thousand carefully curated and deduplicated examples routinely outperform tens of thousands of scraped ones, because the scraped set teaches inconsistency. What matters is that the examples come from real failures, follow one schema, and are scored for correctness before they go in.

Open-weight families, mainly Llama, Mistral, Qwen, Gemma, and Phi, from around 7B up to 70B. Parameter-efficient tuning is what makes that range practical: LoRA for light task adapters, and 4-bit QLoRA using NF4 quantisation, double quantisation, and paged optimizers when the model is too large to train in full precision. We benchmark several candidates against your eval instead of defaulting to a favourite.

At sustained volume on a narrow task, generally yes, and often by around an order of magnitude per token once you account for a small tuned model on your own GPUs instead of a large rented one. At low or spiky volume the arithmetic reverses, because you pay for idle capacity and for the engineering to run it. We model this against your actual request pattern before recommending either way.

It can, and that failure is called catastrophic forgetting. A model trained hard on one narrow task can lose general reasoning and instruction following while scoring well on the thing you measured. We guard against it by keeping general-capability slices in the regression suite, so a candidate that gained on your task and lost elsewhere fails the gate before it ships.

Yes, and for a lot of enterprises that is the whole reason to fine-tune an open-weight model. Both training and serving can run on hardware inside your own data centre, including fully air-gapped with no outbound connectivity, so prompts containing regulated or commercially sensitive data never cross your network boundary. That is usually the deciding factor under HIPAA, PCI, or data-residency obligations, and in defence, critical infrastructure, and any setting where the input itself is the intellectual property. The alternative we also support is your own VPC on AWS, GCP, or Azure, which satisfies most residency requirements while leaving operations to your cloud team.

Yes. Training runs in your own cloud account or on your hardware, and because the base models are open weight there is no requirement to send examples to a hosted tuning service that may retain them. The dataset, the weights, the recipe, and the evaluation harness are all yours at the end, so nothing about the result depends on continued access to us or to a model vendor.

With a harness built before training. We agree the pass criteria up front, hold back a slice of data that training never sees, run deterministic checks first and a judge calibrated against human-labelled examples second, and report the comparison against your current baseline. A checkpoint that does not clear the bar does not ship, and you get the evidence either way.

Foundational pre-training research is not what we do, and neither is chasing benchmark results for their own sake. Our work starts where a model has to become working software. Annotation-only projects go to Payana, our managed data-annotation service.

Ready to scope this stack? Brief the Baaz squad or browse more services.