Fine-tune vs In-Context Learning: The Decision Framework That Actually Works
Part 13, the final post in the AI Engineer Series. The decision between fine-tuning and in-context learning is almost always less obvious than it seems. LoRA, QLoRA, DPO, the eval-driven framework, and the signs that mean "fine-tune now" versus "your prompt has more headroom."
The question almost everyone gets wrong first
You have a feature that is not quite good enough. The accuracy is at 78%, you need 90%, and someone on the team suggests fine-tuning. Before you spend the next six weeks setting up training infrastructure, ask the harder question: have you exhausted what your prompt can do?
Most teams reach for fine-tuning when they should be doing prompt engineering, better retrieval, better examples, or a different model. Fine-tuning is powerful, expensive, and slow to iterate on. In-context learning is fast, flexible, and often gets you most of the way there. The skill is knowing when each one wins, and the answer is almost always less obvious than it seems.
This is Part 13, the final post in the AI Engineer Series. We will cover the decision framework, the techniques (LoRA, QLoRA, DPO), and the eval-driven process that tells you which one is actually working.
What in-context learning can do (and where it stops)
In-context learning (ICL) is everything you do at inference time to steer the model: system prompts, few-shot examples, retrieval-augmented context, chain-of-thought scaffolding, structured output schemas. It is by far the most common technique because it requires no training pipeline, iterates in seconds, and works on any model the same way.
ICL handles the following extremely well: format adherence, style and tone, basic reasoning patterns, integrating fresh knowledge from retrieval, multi-step task decomposition, and structured outputs. If your problem is "the model can do this, I just need to make sure it does this consistently and in the right format," ICL is the right tool.
Where ICL hits a wall: deeply specialized domains where the model lacks foundational knowledge, very long instructions that exceed effective context (yes, models technically support 200K+ tokens, but quality on instruction-following degrades long before that), tasks where the desired behavior contradicts what the model wants to do by default, and high-volume serving where the prompt tokens become a meaningful cost per request.
What fine-tuning can do
Fine-tuning updates the model's weights to bake in behavior that you would otherwise need a long prompt to elicit. Done well, it gives you:
- Substantially shorter prompts (your instructions live in the weights, not the context)
- Stronger and more consistent format adherence
- Domain-specific knowledge and vocabulary that the base model lacks
- Lower inference cost per request, because you can use a smaller model fine-tuned on your task instead of a larger general one
- Behavior that resists prompt-level overrides, which can be a feature or a bug depending on context
What fine-tuning cannot do: replace good evals, fix a bad eval set, teach the model facts that change over time (use retrieval), or rescue a fundamentally underpowered base model. Fine-tuning a small base model on a hard reasoning task usually produces a model that fails fluently in your domain, which is worse than failing obviously in a generic way.
LoRA and QLoRA, the techniques you will actually use
Full fine-tuning updates every weight in the model. It is expensive in compute, memory, and storage, and it produces a separate full-sized model per fine-tune. For most teams, full fine-tuning is overkill.
LoRA (Low-Rank Adaptation) instead trains small adapter matrices alongside the frozen base model. The adapter for a 7B model might be 20-50 MB instead of 14 GB for the full model. LoRA training uses dramatically less GPU memory (typically 4-8x less) and produces tiny artifacts that are easy to manage. Quality is usually within 1-2% of full fine-tuning on most tasks, and the gap closes further with larger ranks or longer training.
QLoRA goes further: it quantizes the frozen base model to 4-bit (Part 10) while training LoRA adapters in higher precision. This lets you fine-tune 70B models on a single 24GB consumer GPU, which was unthinkable a few years ago. QLoRA is the default starting point for serious open-weight fine-tuning today.
For most practical purposes, LoRA or QLoRA with rank 8-16 on the attention projections, trained for 3-5 epochs on a focused dataset, is the default recipe. The Axolotl, Unsloth, and Hugging Face PEFT libraries make this straightforward; you do not need to write CUDA kernels.
DPO and preference tuning, the other major lever
Supervised fine-tuning (SFT) teaches the model to imitate good outputs. Preference tuning teaches the model that certain outputs are preferred over others. The two techniques solve different problems and often compose: SFT to get to "reasonable," then preference tuning to refine the last 10%.
RLHF was the original technique for preference tuning. It is complex, requires a reward model and a PPO pipeline, and is genuinely hard to make work. DPO (Direct Preference Optimization) reformulated the problem as a simple supervised loss over preference pairs (chosen vs rejected response), eliminating the reward model and the reinforcement learning loop. It is dramatically easier to run.
You need preference data: pairs of (prompt, chosen_output, rejected_output). The chosen ones come from your good examples; the rejected ones from earlier model versions, weaker models, or deliberate negatives. A few thousand high-quality pairs is enough to see meaningful behavior shifts. Newer variants like IPO and KTO further refine the loss; for most teams DPO is fine.
The 2024 wave of strong open-weight models (Llama 3, Mistral, Qwen) all use SFT followed by DPO in their fine-tuning recipes. The same recipe works for your tasks. If you have a few thousand human-graded examples from your evals (Part 4), you already have the seed of a preference dataset.
The eval-driven decision framework
Here is the order of operations that actually works.
Start with the eval suite. From Part 4, you have a stratified eval set with deterministic checks, reference metrics, and LLM-as-judge calibrated to humans. If you do not have this, build it before doing anything else. Optimizing without measurement gets you nowhere expensive.
Run the baseline on your current best model and prompt. Note the numbers per stratum (happy path, edge cases, adversarial examples). Pay particular attention to which strata are failing.
Exhaust prompt engineering first. Better instructions, better few-shot examples, better output schemas, better decomposition into smaller LLM calls. This usually closes 50% of the gap between baseline and target in a week. It is fast, cheap, and reversible.
Try a stronger base model. If your numbers are still below target, swap to a more capable model and re-evaluate. Sometimes the right answer is just "use a better model and accept the cost increase," and routing (Part 11) lets you confine that cost to traffic that actually needs it.
Now consider fine-tuning. If you have plateaued at a quality level with the strongest model available, and the failures are concentrated in a specific domain or behavior pattern, fine-tuning becomes attractive. Build a dataset of high-quality examples (typically 1K-10K examples), fine-tune with LoRA, and run the same eval suite. If quality improves on the strata that were failing and doesn't regress elsewhere, you have a winner.
Add preference tuning last. If SFT closes most of the gap but specific subtle behaviors (tone, safety, refusal patterns) are still wrong, DPO with carefully constructed preference pairs is the right tool. Do not start here.
The signs that fine-tuning is the right answer
Specific signals that fine-tuning will actually help, not just feel productive:
- Your prompt is over 2,000 tokens of instructions and you are still seeing inconsistent format adherence
- You have 10,000+ high-quality examples of the desired behavior, with good coverage across edge cases
- Your inference cost is dominated by prompt tokens that could be moved into weights
- The base model lacks vocabulary or knowledge specific to your domain
- Your evals show the model getting the structure right but the substance wrong in domain-specific ways
Signs that fine-tuning is the wrong answer (and you should stay with ICL):
- Your eval set is small or contaminated (more eval work needed)
- The task is fundamentally about combining current information with reasoning (use retrieval)
- The desired behavior changes every few weeks (retraining cycle is too slow)
- You haven't tried the bigger model yet
- Your engineering team has never deployed a fine-tuned model in production before (start with the easier wins first)
The honest scorecard
In-context learning gets you to 80-90% of the way for most production use cases. Fine-tuning is the right tool for the remaining 10-20%, in specific situations, after you have eval infrastructure in place.
The teams that win at fine-tuning are not the ones that fine-tune first. They are the ones that built rigorous evals, exhausted prompt engineering, tried better models with routing, and then fine-tuned with clear hypotheses about what would improve and how they would know. The success rate on that path is high. The success rate on "let's just fine-tune to see if it helps" is roughly zero.
And that is the series
Across thirteen posts, we covered the engineering disciplines that turn LLM features from demos into reliable production systems: the harness around your calls, retrieval and context, structured outputs, evals, observability, latency and throughput, caching, cost attribution, KV cache, inference acceleration, routing, agent safety, and now training-time techniques.
None of this is glamorous. None of it is the part that goes in the keynote. All of it is the part that decides whether your AI feature is something users rely on, or something they avoid. The discipline compounds. Teams that take all thirteen seriously ship faster, cost less, and break less. Teams that take none of them seriously have a different career path.
Build the harness. Measure everything. Stay practical. The frontier moves fast, the engineering fundamentals do not.
Previous in the series
- Part 1: Harness Engineering
- Part 2: Context Engineering & Retrieval Quality
- Part 3: Structured Outputs and Fallback Chains
- Part 4: Evals
- Part 5: LLM Observability
- Part 6: Latency and Throughput Engineering
- Part 7: Prompt Caching vs Semantic Caching
- Part 8: Cost Attribution Per Feature
- Part 9: The KV Cache
- Part 10: Speculative Decoding and Quantization
- Part 11: Model Routing and Graceful Fallback
- Part 12: Agent Guardrails and Loop Budgets
References
- LoRA: Low-Rank Adaptation of Large Language Models by Hu et al. The paper behind most efficient fine-tuning today.
- QLoRA: Efficient Finetuning of Quantized LLMs by Dettmers et al. The quantization-plus-LoRA combination that put 70B fine-tuning on consumer hardware.
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model by Rafailov et al. The DPO paper that made preference tuning practical.
- Hugging Face PEFT. The standard library for LoRA, QLoRA, and other parameter-efficient fine-tuning methods.
- Fine-tuning FAQ by Hamel Husain. The most practical writing on when to fine-tune and how to do it well.
- Patterns for Fine-tuning by Eugene Yan. Overview of the techniques and how they compose.