Results were a mixed, which is what made it interesting. Given the right crib, the Bombe breaks 22 of 25 messages; ciphertext-only against 10 plugs, 0 of 25, same wall Bletchley hit. As a judge of "is this decryption German?" Jev is excellent: AUC 0.9999, 159/160 right decisions. As a crib selector it has no skill at all; its ranking is no better than the fixed list.
Against XGBoost an classical judges, it ties on the same kind of data. But trained on synthetic traffic and tested on real intercepts, and XGBoost drops to 53% correc versus zero-shot Jev holds at 97%. Feel free to try it here: https://enigma-jev.vercel.app/ (requires a jev typesafe api key)