← Back to Blog

Llama 3 8B Fine-Tuning: My Legal Doc Classifier Project Was a Reality Check

57 Reads
Llama 3 8B Fine-Tuning: My Legal Doc Classifier Project Was a Reality Check

Look, everyone's been buzzing about Llama 3 8B. And for good reason, right? Meta pushed out a beast. Faster, sharper, better at following instructions right out of the box than almost anything else at its size. I was hyped. We all were. So when Dr. Anya Sharma from our legal tech division asked if I could get it to classify very specific legal document types for our internal 'Project Chimera,' I told her, "Yeah, no sweat. Give me a week."

Turns out, sweat was very much involved.

Our task was narrow: precisely distinguish between 17 categories of legal filings—think specific patent applications, trademark registration forms, initial discovery requests, and some really niche compliance documents. We already had a custom BERT model doing okay, hitting about 87% F1-score. But it was brittle. Small changes in phrasing or formatting threw it off. The hope was Llama 3, fine-tuned, could be more robust, more 'intelligent.'

I started simple. Grabbed 320 hand-labeled examples from our existing dataset—about 18-20 per class. My first pass used QLoRA on a rented A100 from RunPod. Four hours of training time, minimal hyperparameter tuning. My office AC unit decided that was the perfect day to kick the bucket, too. Ninety-degree heat, sweat dripping, and the monitor showing a model that... well, it wasn't terrible. It wasn't great either. Accuracy hovered around 70%. It kept confusing trademark applications with patent submissions. Big problem.

My initial assumption was simple: more positive examples. So I went back to the data engineers. "Give me everything you've got!" Over the next 9 days, we amassed almost 1,800 total examples. I trained again. Better, sure, but only nudging 78%. Still too many false positives. The model was aggressively over-generalizing. It was like teaching a toddler to identify dogs by only showing them golden retrievers, then being surprised when they called every four-legged animal a "doggo."

# This was my initial LoRA config. Too simple, I found. from peft import LoraConfig, get_peft_model lora_config = LoraConfig( r=8, lora_alpha=16, target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM" ) # ... later applied to the model

It took me a solid two weeks of iterating, debugging, and drinking way too much lukewarm instant coffee to figure out the negative examples were the real bottleneck. Not just 'other' documents, but specific, carefully constructed non-examples that were structurally similar but semantically distinct. For every five positive examples of a patent application, I realized I needed two examples that looked like patent applications but were definitively not. My dog, Buster, wasn't helping, barking at every squirrel outside my window while I was trying to manually curate these tricky non-matches.

After painstakingly adding nearly 500 such carefully selected 'negative' data points, retraining the Llama 3 8B model finally pushed our F1-score to 91.3% for the critical categories, with a p95 inference latency of 180ms on our deployed endpoint. That's a huge win against the old BERT model.

Here’s my actual, slightly controversial take: for many routine classification tasks, particularly those where you're just trying to map text to a known label, you might actually be better off perfecting a few-shot prompt with the Llama 3 70B model than wrestling with fine-tuning the 8B model. The 70B model’s raw reasoning power, even with only a few examples in the prompt, can often exceed a hastily or imperfectly fine-tuned smaller model. Fine-tuning 8B still demands meticulous data prep; it's not a shortcut around bad data.

Final Thoughts

Llama 3 8B is fantastic. No doubt. But the idea that a smaller, fine-tuned model instantly solves all your domain-specific problems without significant data curation is just wishful thinking. The underlying principles of good machine learning haven't changed: garbage in, garbage out. For Project Chimera, it worked, eventually. But it was a hard-fought battle, proving that even with a state-of-the-art base model, your data is still king. Don't underestimate the negative examples.