- September 2026 (1)
- August 2026 (4)
- July 2026 (8)
- June 2026 (7)
- May 2026 (2)
- April 2026 (11)
- March 2026 (3)
- February 2026 (4)
- January 2026 (4)
- December 2025 (1)
- November 2025 (3)
- October 2025 (9)
- September 2025 (3)
- August 2025 (5)
- July 2025 (1)
- June 2025 (2)
- May 2025 (3)
- April 2025 (2)
- March 2025 (7)
- February 2025 (10)
- January 2025 (6)
- December 2024 (7)
- September 2024 (1)
- August 2024 (2)
- July 2024 (2)
- May 2024 (2)
- April 2024 (2)
- February 2024 (2)
- April 2023 (1)
- March 2023 (2)
- September 2022 (1)
- February 2022 (1)
- November 2021 (1)
- March 2021 (1)
- February 2021 (2)
- August 2019 (1)
- November 2018 (1)
- May 2017 (1)
- December 2016 (1)
- April 2016 (1)
- August 2015 (1)
- December 2014 (1)
- August 2014 (1)
- March 2014 (1)
- December 2013 (1)
- October 2013 (3)
- September 2013 (4)
- August 2013 (2)
- July 2013 (1)
- June 2013 (1)
- February 2013 (1)
- October 2012 (1)
- June 2012 (1)
- May 2012 (1)
- April 2012 (1)
- February 2012 (1)
- October 2011 (1)
- June 2011 (1)
- May 2011 (1)
- April 2011 (1)
- March 2011 (1)
- February 2011 (1)
- January 2011 (1)
- December 2010 (3)
- November 2010 (1)
- October 2010 (1)
- September 2010 (1)
- August 2010 (1)
- July 2010 (1)
- May 2010 (3)
- April 2010 (1)
- March 2010 (2)
- February 2010 (3)
- January 2010 (4)
- December 2009 (2)
- November 2009 (5)
- October 2009 (2)
- September 2009 (2)
- August 2009 (3)
- July 2009 (1)
- May 2009 (1)
- April 2009 (1)
- March 2009 (5)
- February 2009 (5)
- January 2009 (5)
- December 2008 (3)
- November 2008 (7)
- October 2008 (4)
- September 2008 (2)
- August 2008 (1)
- July 2008 (1)
- June 2008 (1)
- May 2008 (1)
- April 2008 (1)
- January 2008 (4)
- December 2007 (3)
- March 2007 (3)
- February 2007 (1)
- January 2007 (2)
- December 2006 (4)
- November 2006 (18)
- AI (99)
- Python (78)
- TIL deep dives (77)
- LLM from scratch (48)
- Resolver One (34)
- PyTorch (23)
- TIL (22)
- Blogkeeping (20)
- PythonAnywhere (17)
- Linux (16)
- Startups (15)
- Hugging Face (14)
- Gadgets (13)
- NSLU2 offsite backup project (13)
- Funny (11)
- Musings (11)
- Finance (10)
- Fine-tuning LLMs (10)
- JAX (10)
- C (9)
- Website design (9)
- Personal (8)
- Robotics (8)
- 3D (5)
- Quick links (5)
- Rants (5)
- Cryptography (4)
- GPT-2 mysteries (4)
- JavaScript (4)
- Music (4)
- Oddities (4)
- Talks (4)
- Dirigible (3)
- Eee (3)
- Memes (3)
- Politics (3)
- Django (2)
- GPU Computing (2)
- LaTeX (2)
- MathML (2)
- Microprojects (2)
- OLPC XO (2)
- Retro Language Models (2)
- Space (2)
- VoIP (2)
- Copyright (1)
- Golang (1)
- MoE models (1)
- poppy the training box (1)
- Raspberry Pi (1)
- Software development tools (1)
- Agile Abstractions
- antirez
- Astral Codex Ten
- :: (Bloggable a) => a -> IO ()
- David Friedman's Substack
- Econ & Energy
- Entrepreneurial Geekiness
- For some value of "Magic"
- Hackaday
- kaleidic.ai newsletter
- Knowing.NET
- Language Log
- Millennium Hand
- ntoll.org
- Obey the Testing Goat!
- One Useful Thing
- PK
- PythonAnywhere News
- Simon Willison's Weblog
- Societive
- Software Deviser
- Some opinions, held with varying degrees of certainty
- tartley.com
- the singularity is nearer
- Theia Vogel's website
poppy the training box, part 1: the beginnings
For a while I've been planning to put together a separate machine for local LLM
training. Until now, I've been using my desktop PC, perry. I have an RTX 3090
installed, and can get useful training runs done (most recently,
a 163M-parameter GPT-2 small style LLM in JAX),
but there are a couple of problems.
perryis my daily driver. If he's doing a training run, then everything is just a little bit sluggish as CPU and GPU alike are busy.- Although I don't play games often, it's annoying to have the option ruled out for days at a time.
- While the GPU is busy with a training run, I can't do other experiments in parallel -- for example, to scope out what the next step might be.
And relatedly to all of those: the two-day limit to the training runs I've been doing is something I set
because that's the maximum amount of time I'm willing to have perry tied up. It
would be really interesting to try longer training runs!
I also have longer-term plans; a multi-GPU box would be interesting to put together -- not just to have more power locally, but so that I could test larger-scale cloud multi-GPU training runs before starting to pay for expensive machines. US$15.92 an hour to rent a machine isn't a lot of money, but it adds up, especially if you're spending it while debugging parallelism issues.
And finally, I've always been interested in putting together a custom water-cooling loop in a PC. I've been building my own machines since 1995 or so, but never got round to that side of things. It sounds fun!
But despite all of those future plans, this is a fairly normal machine-building post -- how I repurposed an old PC, plugged in a second-hand RTX 3090 from eBay, tested it all, accidentally trained an LLM for 11 days, and almost cooked a CPU.
Over time, I expect to be posting more -- and more interesting -- build details. Let's think of this as establishing the baseline.