I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.
Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.
whythismatters•37m ago
https://github.com/jessewaites/antiquity
yannis•25m ago