Skip to content

Comment on Training NanoGPT on My Journalparent

Comments

True, true. But I already get somewhat coherent sentences that are personally enjoyable and valuable.

The main problem is that 90% of the output is garbage: broken formatting, abruptly ending sentences, mixed languages, occasional nonsensical sentences like `your "two" type,ir zwe "G together"`. If I can produce something that looks like normal sentences even if they're not particularily meaningful, with ~90% reliability, that's all I need for my purposes, and I think the current bottleneck is actually the very poor dataset of my very heterogeneous notes.

I'm thinking about it like how some people love their pets or toddlers even though they struggle with the most basic tasks. The personal connection makes up for it :)

Yeah, if you only want coherent language, 1B parameters is more or less good enough. If you also want general context awareness and some basic world knowledge, you will need to up it another order of magnitude. Either way, this is nothing that you'd run on dated hardware. Smaller transformers can be useful for very specific task, but autoregressive LMs are just insane when it comes to hardware requirements.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.