We are considering a combination of following -
1. How much can it do on its own with minimal to no human oversight (can it write a test case, automate, setup testing pre-reqs, run the tests, generate report etc) and how much is the approval percent when human tester reviews agent's work (how less do I have to correct its output)
2. How faster are the releases now, since AI is known for speed if we are doing faster releases or more releases that translates to better ROI
3. What is the count of bugs found / avoided using AI (if production is cleaner than before or not)
4. Cost of AI while trying to do all of the above.
----------------------------------------
One major problem we have faced is the non deterministic and hallucinating nature of AI, we have tried to reduce it by having bash / python scripts in skills / rules so that some level of deterministic output is guaranteed - so far worked but again depends on use cases as well.