- August 2026 (1)
- July 2026 (8)
- June 2026 (7)
- May 2026 (2)
- April 2026 (11)
- March 2026 (3)
- February 2026 (4)
- January 2026 (4)
- December 2025 (1)
- November 2025 (3)
- October 2025 (9)
- September 2025 (3)
- August 2025 (5)
- July 2025 (1)
- June 2025 (2)
- May 2025 (3)
- April 2025 (2)
- March 2025 (7)
- February 2025 (10)
- January 2025 (6)
- December 2024 (7)
- September 2024 (1)
- August 2024 (2)
- July 2024 (2)
- May 2024 (2)
- April 2024 (2)
- February 2024 (2)
- April 2023 (1)
- March 2023 (2)
- September 2022 (1)
- February 2022 (1)
- November 2021 (1)
- March 2021 (1)
- February 2021 (2)
- August 2019 (1)
- November 2018 (1)
- May 2017 (1)
- December 2016 (1)
- April 2016 (1)
- August 2015 (1)
- December 2014 (1)
- August 2014 (1)
- March 2014 (1)
- December 2013 (1)
- October 2013 (3)
- September 2013 (4)
- August 2013 (2)
- July 2013 (1)
- June 2013 (1)
- February 2013 (1)
- October 2012 (1)
- June 2012 (1)
- May 2012 (1)
- April 2012 (1)
- February 2012 (1)
- October 2011 (1)
- June 2011 (1)
- May 2011 (1)
- April 2011 (1)
- March 2011 (1)
- February 2011 (1)
- January 2011 (1)
- December 2010 (3)
- November 2010 (1)
- October 2010 (1)
- September 2010 (1)
- August 2010 (1)
- July 2010 (1)
- May 2010 (3)
- April 2010 (1)
- March 2010 (2)
- February 2010 (3)
- January 2010 (4)
- December 2009 (2)
- November 2009 (5)
- October 2009 (2)
- September 2009 (2)
- August 2009 (3)
- July 2009 (1)
- May 2009 (1)
- April 2009 (1)
- March 2009 (5)
- February 2009 (5)
- January 2009 (5)
- December 2008 (3)
- November 2008 (7)
- October 2008 (4)
- September 2008 (2)
- August 2008 (1)
- July 2008 (1)
- June 2008 (1)
- May 2008 (1)
- April 2008 (1)
- January 2008 (4)
- December 2007 (3)
- March 2007 (3)
- February 2007 (1)
- January 2007 (2)
- December 2006 (4)
- November 2006 (18)
- AI (96)
- TIL deep dives (77)
- Python (75)
- LLM from scratch (48)
- Resolver One (34)
- PyTorch (21)
- TIL (21)
- Blogkeeping (19)
- PythonAnywhere (17)
- Linux (16)
- Startups (15)
- Gadgets (13)
- Hugging Face (13)
- NSLU2 offsite backup project (13)
- Funny (11)
- Musings (11)
- Finance (10)
- Fine-tuning LLMs (10)
- C (9)
- JAX (9)
- Personal (8)
- Robotics (8)
- Website design (8)
- 3D (5)
- Quick links (5)
- Rants (5)
- Cryptography (4)
- JavaScript (4)
- Music (4)
- Oddities (4)
- Talks (4)
- Dirigible (3)
- Eee (3)
- GPT-2 mysteries (3)
- Memes (3)
- Politics (3)
- Django (2)
- GPU Computing (2)
- LaTeX (2)
- MathML (2)
- Microprojects (2)
- OLPC XO (2)
- Retro Language Models (2)
- Space (2)
- VoIP (2)
- Copyright (1)
- Golang (1)
- poppy the training box (1)
- Raspberry Pi (1)
- Software development tools (1)
- Agile Abstractions
- antirez
- Astral Codex Ten
- :: (Bloggable a) => a -> IO ()
- David Friedman's Substack
- Econ & Energy
- Entrepreneurial Geekiness
- For some value of "Magic"
- Hackaday
- kaleidic.ai newsletter
- Knowing.NET
- Language Log
- Millennium Hand
- ntoll.org
- Obey the Testing Goat!
- One Useful Thing
- PK
- PythonAnywhere News
- Simon Willison's Weblog
- Societive
- Software Deviser
- Some opinions, held with varying degrees of certainty
- tartley.com
- the singularity is nearer
- Theia Vogel's website
Writing an LLM from scratch, part 26 -- evaluating the fine-tuned model
This post is on the second half of chapter 7 of Sebastian Raschka's book "Build a Large Language Model (from Scratch)". In the last post I covered the part of the chapter that covers instruction fine-tuning; this time round, we evaluate our model -- particularly interestingly, we try using another, smarter, model to judge how good its responses are.
Once again, Raschka's explanation in this section is very clear, and there's not that much that was conceptually new to me, so I don't have that many notes -- in fact, this post is probably the shortest one in my series so far!
Generating the test set responses
Unusually, when at the start of section 7.7 we generate some sample responses for the instructions in our test set, I got exactly the same results as in the book. For once, I guess, everything that uses randomness was happening in the same order as it did when Raschka ran it on his machine.
The next step was to generate a file with all of the responses to all of the test instructions, which took 18.9 seconds on my RTX 3090 (compared to a minute on an A100, per the book -- that's quite surprising!)
Once that was done, it was time to install Ollama so that I could use the Llama 3 model to evaluate my own.
Ollama
I've never used Ollama before -- when playing with other people's models, I've always used Hugging Face's Transformers library.
It's a neat package, though. It wraps llama.cpp, which is a pure C/C++ inference
framework (with CUDA support), and makes it easy to download and run models that
have been packaged for it. Being written in C, I would imagine that it's faster than
PyTorch/Transformers -- though, being inference-only, it's less useful if you're planning to do things
like training or fine-tuning the models.
My desktop is running a fairly customised install of Arch Linux, and I didn't want to
use the default install procedure (which puts it into your system-wide /bin and /lib
directories). But it turns out that it's a very well-packaged app,
and you don't need to do that.
Using the manual install instructions for Linux,
I just created a new directory ~/Dev/ollama, and then cded there and downloaded it:
wget https://ollama.com/download/ollama-linux-amd64.tgz
It was about 1.75 GiB. I then untarred it:
tar xf ollama-linux-amd64.tgz
...and then I could run commands with full paths, for example:
~/Dev/ollama/bin/ollama serve
...to start up the server, or
~/Dev/ollama/bin/ollama run llama3
...to start a session.
Neat! It's always good to see pre-built binary packages that have no issues with their install location.
Actually running the evaluation
The next step was to throw all of the generated test responses (and their associated targets) at Llama 3 and see what it thought about how close they were.
Again, this all worked without trouble. I noted that the responses I was getting from Llama 3
were not the same as the ones in the book -- Raschka notes that
Ollama is non-deterministic, so there's no surprise there (though it does make me
wonder why it accepts a seed parameter in the API call).
When I got on to the final eval, where you run the test results through Llama 3 and ask it to rate them compared to the target outputs, it took 11 seconds to run, and I got an average score of 48.95 / 100, which is close enough to the 50.32 that appears in the book. 1 I'd run an eval on my model, using a smarter model to judge its responses!
Somewhat surprisingly, that number was stable over multiple runs. So perhaps there
is some level of determinism in Ollama now that wasn't present when the book was written,
and the seed (eg. 123) is of value. Or perhaps Raschka's comment about it being non-deterministic
was more of a "between machines" thing rather than for multiple runs on the same machine
-- though then I'm not sure why he suggests re-running it for multiple results.
Anyway -- that was it! Eval done. And, to my amazement, that was the end of the chapter -- and almost the end of the book. We've built an LLM from scratch, fine-tuned it, and evaluated it by using a smarter model to judge how well it was following instructions.
This is the end...
...or at least the end of the beginning.
Having run the evaluation, I've reached the end of the main part of "Build a Large Language Model (from Scratch)". But I don't think I've reached the end of this project, there's still more to do (not least working through the appendices).
So, coming up next: a post summarising what I've got through so far in this series, and what the next steps are to wrap it up.
Here's a link to the next post in this series.
-
I also got 110 out of 110 scores -- that is, every response from Llama 3 was parseable as an integer. That actually kind of surprised me! Models like to be chatty and helpful. But looking into it, the famous X post by Riley Goodside where he had to "threaten" Bard to stop it from saying "Sure, no problem! Here's your JSON" was almost two years ago. ↩