For those running open-source AI infrastructure in production:
- What are you running, for what workload, and would you recommend?
- Do you operate yourself versus consume as a managed service?
- Have you replaced or abandoned any tools because they were too difficult or expensive to operate?
- What problems only appeared after moving beyond the prototype stage?
- Anything that you would do differently if rebuilding the stack today?
Thanks