de en es fr nl pl pt sv zh

training

Collections: Pre-Modern Armies for Worldbuilders, Part IVb: Cohesion

Bret Devereaux

This is the second half of the fourth part (I, IIa, IIb, III, IVa) of our honestly-who-knows-how-many part series laying out some general guidelines for how pre-modern armies are organized. Last week we took a look at how the leadership of these armies was drawn from the existing elite classes of...

It’s always the learning rates

Accidental Factors

Pre-training any kind of good LLM is very, very expensive. Thankfully, we have scaling laws. Lilian Weng of Thinky writes: Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss decreases predictably as we scale up...

FactWorld

Accidental Factors

When we started building LLMs, we mostly focused on them knowing things. They had information encoded in their weights, and they could spit it out when given sufficient prompts. But an agent doesn’t just need to know things; it needs to combine several kinds of knowledge. A lot of that is still...

Somehow, more on distillation

Accidental Factors

The capabilities in a large language model emerge, mysteriously, from the training data. Everyone agrees that you start with a big pile of data, add some compute, and at the end you can vibe code. Opinions differ on what that pile of data should look like. Microsoft AI recently released an...

2022 Holiday Training Sale

Chris Sanders

Once a year, all of my training courses go on sale. This year, that sale starts now, on November 25th, and runs until midnight (ET) on November 29th. All of my courses are 20-25% off, and the website prices currently reflect that discount. You can register for a course at...