hey that's awesome! yeah the eval showed first pass was only ~65% real decisions. The fix that stuck was an entry has to name a real file it touches or it gets dropped. A code decision names code.
I agree agents don't always self-talk decisions, that's why we distill the whole transcript after the fact instead of asking them to log anything. Your baselining idea is good!
Comments
hey that's awesome! yeah the eval showed first pass was only ~65% real decisions. The fix that stuck was an entry has to name a real file it touches or it gets dropped. A code decision names code.
I agree agents don't always self-talk decisions, that's why we distill the whole transcript after the fact instead of asking them to log anything. Your baselining idea is good!