An agent reviewing its own output has the same blind spots, the same training-data gaps, and a strong bias toward declaring its work "good". But even having a different model review the code is flawed and not good enough, especially for those whose code can't be wrong, like in regulated industries.
Comments
An agent reviewing its own output has the same blind spots, the same training-data gaps, and a strong bias toward declaring its work "good". But even having a different model review the code is flawed and not good enough, especially for those whose code can't be wrong, like in regulated industries.