Cool that you guys created this, but interpreting the results: why wouldn't I just use Claude instead of an AI SRE tool?
Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?
rafiyashaheen11•57m ago
How much time did it take to build this? Loved it! What problem does this solve?
emrahsamdan•26m ago
Running the test takes around 5-6 hours per vendor depending on how much wrestle it requires (not every system is ideal for ai agents - including ourselves) but the actual work was to come up with the real world scenarios and what would be the expected RCA and remediation for this. Our own SRE team worked for 1.5 weeks for that.
esafak•16m ago
Emrah, it seems that EDX does not offer any edge over GCX, when paired with an agent like Claude?
nikhilunni•1h ago
Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?