Skip to content

Ask HN: What are your go-to "test" questions when evaluating a new LLM?

7 pointsjohntiger13 comments
On HN

Do you have a go-to question (or several) to check if an LLM knows its stuff? For me, I ask a simple question:

"What is Operation Konrad III"

which most LLMs fail due to the (relative) obscurity of the event.

Comments

Not really scientific or anything but I tend to give it the task: „Write a simple http server in Go that saves all requests into a SQLite database.“

What I am looking for is:

- did it forget to import the SQLite driver?

- is it doing weird SQL shenanigans like selecting MAX(id) to obtain the next potential id?

- is the code rather simple or over-engineered?

update: Most LLMs produce a decent answer, however it you increase the difficulty a little bit by asking it "Write a simple and CGo free http server in Go ...", most LLMs get the sql driver wrong (except for gpt-4-1106-preview)

I give it a large block of code and see if it can find the bug. Amusingly, GPT sometimes passes it with flying colors (finding the bugs I didn't see and seeing unused imports) but at other times it just flat out fails to see anything.

I ask it about creating a conversation in Polish with English translations about an encounter between two neighbours walking their dogs.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.