OpenInference and OpenTelemetry have a schema for agent sessions.
I finally discovered agentsview and ctxrs/ctx for indexing and searching agent session logs. Perhaps such an index is also useful for agent optimization
Nonce
Accuracy is sensitive to
nonce and temperature without variance in the model hyperparameters
Is this a convex optimization or non-convex optimization problem?
- Neither, frugon sidesteps that issue by having a small finite candidate selection that it evaluates completely. No search needed because the problem is discrete and small.
OpenInference and OpenTelemetry have a schema for agent sessions.
- Good to know for when I look into it.
Accuracy is sensitive to nonce and temperature without variance in the model hyperparameters
- Yes, this is a current weakness in frugon...Hmm, I see, so I'd essentially need to establish a baseline (with whatever the default temp is) + nonce, then compare both spreads. Noted.
Good thread. Give frugon a try on your own logs, and tell me what you think.
As well. Someday costed opcodes and sharded redundancy for all of this;
No search needed because the problem is discrete and small.
If an algorithm selects according to rank, there's a cost/score/fitness/survival/error function (that assumes or imposes a distance metric to make a metric space).
Awhile back I had to match exterior paint to printed paint swatches and I'm stereotypically not good at colors, so finally I arrived at just A/B; this one or this one and discard.
As well. Someday costed opcodes and sharded redundancy for all of this;
Hmm, noted.
If an algorithm selects according to rank, there's a cost/score/fitness/survival/error function.
- Yes, what I'm saying is that frugon has a small enough candidate pool that it can all be evaluated at the same time. It doesn't need to search through it.
Awhile back I had to match exterior paint to printed paint swatches and I'm stereotypically not good at colors, so finally I arrived at just A/B; this one or this one and discard.
- This is exactly the shape the judge prompt takes in frugon.
Comments
Is this a convex optimization or non-convex optimization problem?
Re: pytest-evals, mcpbr, agentevals, foundry-toolkit; and devtools-mcp; https://news.ycombinator.com/item?id=48532642
OpenInference and OpenTelemetry have a schema for agent sessions.
I finally discovered agentsview and ctxrs/ctx for indexing and searching agent session logs. Perhaps such an index is also useful for agent optimization
Accuracy is sensitive to nonce and temperature without variance in the model hyperparameters
Down the rabbit hole we go.
- Neither, frugon sidesteps that issue by having a small finite candidate selection that it evaluates completely. No search needed because the problem is discrete and small.
- Good to know for when I look into it.
- Yes, this is a current weakness in frugon...Hmm, I see, so I'd essentially need to establish a baseline (with whatever the default temp is) + nonce, then compare both spreads. Noted.
Good thread. Give frugon a try on your own logs, and tell me what you think.
As well. Someday costed opcodes and sharded redundancy for all of this;
If an algorithm selects according to rank, there's a cost/score/fitness/survival/error function (that assumes or imposes a distance metric to make a metric space).
Awhile back I had to match exterior paint to printed paint swatches and I'm stereotypically not good at colors, so finally I arrived at just A/B; this one or this one and discard.
Hmm, noted.
- Yes, what I'm saying is that frugon has a small enough candidate pool that it can all be evaluated at the same time. It doesn't need to search through it.
- This is exactly the shape the judge prompt takes in frugon.
See: src/frugon/measure.py
Then search: JUDGE_PROMPT_TEMPLATE