We can do a mix of general use (as in user stories) plus a few academic benchmarks.
Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.
So you have any specific ideas?
Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability.
should test the harness capability rather than model's knowledge/capability.
Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class.
Comments
Interesting idea. If you can manage the infra, I can put together a replicable test suite.
I can manage the infra, have a lot of experience in that area. A benchmark with problems coming from multiple sources and backgrounds would be ideal
We can do a mix of general use (as in user stories) plus a few academic benchmarks.
So you have any specific ideas?
You can hmu at iam@thechris.in
Thanks, I will reach out. I have also posted for contributors on localllama https://www.reddit.com/r/LocalLLaMA/comments/1vg40w8/anyone_...
Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.
Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability.
Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class.