I have a book in PDF that I wanted to read on my Boox, the ebook reader for non-Kindle users. It's perfectly capable of viewing PDFs, but we're in the age of intelligence, and settling for a non-native format seemed like a tradeoff I wasn't willing to make.
I thought it was the perfect straightforward task to delegate to my agent. "Please convert this PDF to EPUB" I prompted confidently. Little did I know it was far from it. The pages were all scanned, which is more likely for older books with no native digital versions. The agent pressed on, ignored all the diagrams and outputted a digital book with garbled text. With further direction from me, it was able to identify and extract all diagrams and place them in their correct position. But the entire process dragged on for over two hours, yielding uneven fonts, awkward spacing, missing captions, and miserable formatting throughout.
Yes, I went in not knowing there's no linear path to converting scanned books to digital versions, and maybe that's on me. But thanks to reinforcement learning, these models are trained to keep digging in their heels and go on an endless quest for a solution. I would've just settled for "Listen, that's not possible, and my best efforts would render something resembling a digital book, but not quite readable by anyone. Would you like to proceed anyway?" Instead, it was a performative circus, lighting tokens on fire.
Where the tokens went
cumulative estimated token spend across the session