https://ai.facebook.com/blog/dynaboard-moving-beyond-accuracy-to-holistic-model-evaluation-in-nlp Last year, Facebook AI released Dynabench, a platform that radically rethinks benchmarking in AI, starting with natural language processing (NLP) models. Going forward, they have now announced a new evaluation-as-a-service platform for comprehensive, standardized evaluations of NLP models called Dynaboard. Dynaboard can perform apples-to-apples comparisons dynamically without common issues from bugs in evaluation code, inconsistencies in filtering test data, backward compatibility, accessibility, and several other reproducibility issues. Dynaboard enables AI researchers to customize a new Dynascore metric based on multiple axes of evaluation, including compute, accuracy, robustness, memory, and fairness. https://ai.facebook.com/blog/dynaboard-moving-beyond-accuracy-to-holistic-model-evaluation-in-nlp Dynascore allows AI researchers to dynamically adjust the default score by placing more or less weight on particular metrics to evaluate performance comprehensively. This capability is an essential feature of Dynascore since every person who uses leaderboards has different preferences and goals. Since launching Dynabench, Facebook AI has collected over 400,000 examples and has released two new, challenging data sets. Facebook AI believes that as the AI community continues to build on its open platform, the field will iteratively and rigorously improve how researchers evaluate models, create data sets, and eventually evolve towards better benchmarks. Dynaboard offers maximum flexibility for users who want to make fine-grained comparisons between…
Read More











