LDB: Large Language Model Debugger via Verifying Runtime Execution Step by Stepgithub.com/FloridSleeves 2 pointspanqueca2 years ago1 commentSaveHideCopy link On HNComments−panquecaOP2yHumanEval Benchmark: 95.1 @ GPT-3.5I wonder if it can be combined with projects like SWE-Agent to build powerful yet opensource coding agents.- https://paperswithcode.com/sota/code-generation-on-humaneval- https://github.com/princeton-nlp/SWE-agent
Comments
HumanEval Benchmark: 95.1 @ GPT-3.5
I wonder if it can be combined with projects like SWE-Agent to build powerful yet opensource coding agents.
- https://paperswithcode.com/sota/code-generation-on-humaneval
- https://github.com/princeton-nlp/SWE-agent