There are a number of AI benchmarks but I feel they aren't practical when it comes to evaluating AI products. Several months back we started working on designing a methodology to independently evaluate AI products. After conducting a good bit of research, we tested the methodology on a few different categories of products and improvised it.
If this is an area of interest to you, may I request you to review this methodology and help us with feedback. A quick blurb on the structure of the methodology - It's a 3-tier framework starting with 6 high level evaluation categories (Quality, Security, Privacy & safety, Use cases & pricing, Sustainability & ecosystem, and Impact & ethics).
The methodology is available at huby.ai/methodology.
Thanks in advance.