Skip to content

Comment on Show HN: Frugon – Find which LLM calls a cheaper model could handle (local, MIT)parent

Comments

minimize cost to converge upon sufficiently low error problem

Is this a convex optimization or non-convex optimization problem?

Evals

Re: pytest-evals, mcpbr, agentevals, foundry-toolkit; and devtools-mcp; https://news.ycombinator.com/item?id=48532642

OpenInference and OpenTelemetry have a schema for agent sessions.

I finally discovered agentsview and ctxrs/ctx for indexing and searching agent session logs. Perhaps such an index is also useful for agent optimization

Nonce

Accuracy is sensitive to nonce and temperature without variance in the model hyperparameters

Down the rabbit hole we go.

Is this a convex optimization or non-convex optimization problem?

- Neither, frugon sidesteps that issue by having a small finite candidate selection that it evaluates completely. No search needed because the problem is discrete and small.

OpenInference and OpenTelemetry have a schema for agent sessions.

- Good to know for when I look into it.

Accuracy is sensitive to nonce and temperature without variance in the model hyperparameters

- Yes, this is a current weakness in frugon...Hmm, I see, so I'd essentially need to establish a baseline (with whatever the default temp is) + nonce, then compare both spreads. Noted.

Good thread. Give frugon a try on your own logs, and tell me what you think.

As well. Someday costed opcodes and sharded redundancy for all of this;

No search needed because the problem is discrete and small.

If an algorithm selects according to rank, there's a cost/score/fitness/survival/error function (that assumes or imposes a distance metric to make a metric space).

Awhile back I had to match exterior paint to printed paint swatches and I'm stereotypically not good at colors, so finally I arrived at just A/B; this one or this one and discard.

As well. Someday costed opcodes and sharded redundancy for all of this;

Hmm, noted.

If an algorithm selects according to rank, there's a cost/score/fitness/survival/error function.

- Yes, what I'm saying is that frugon has a small enough candidate pool that it can all be evaluated at the same time. It doesn't need to search through it.

Awhile back I had to match exterior paint to printed paint swatches and I'm stereotypically not good at colors, so finally I arrived at just A/B; this one or this one and discard.

- This is exactly the shape the judge prompt takes in frugon.

See: src/frugon/measure.py

Then search: JUDGE_PROMPT_TEMPLATE

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.