The most difficult part was detecting what frames in the video are slides, and extracting from those what is useful content to add to the document. I didn't want to just stick a screengrab of the slides in the text, I wanted it to read like an actual book, so it uses an AI model to take from the slides what is useful and ignores what is just repeated in the transcript.
Here is an output sample from an NIH video: https://lecturetobook.com/sample/sleep-book/book.html
Feedback welcome! This is my first side-project that has made it to launch.