I assume that this is "$ spent on search + $ spent on LLM" < budget, but how do you handle the LLM spending more than you would expect on a request? Or is this handled by max_tokens and some form of pricing table? (and if so, how does caching play a role?)
I'm glad your numbers are honest! For a moment I thought, hey, maybe this person's numbers are lying to me... but it turned out they were not so thank you!
conception 1 hours ago [-]
They are honest because they load bear the seam.
reindeer2 5 hours ago [-]
[dead]
Rendered at 07:03:58 GMT+0000 (Coordinated Universal Time) with Vercel.
I assume that this is "$ spent on search + $ spent on LLM" < budget, but how do you handle the LLM spending more than you would expect on a request? Or is this handled by max_tokens and some form of pricing table? (and if so, how does caching play a role?)
[0]: https://www.datamole.ai/
I'm glad your numbers are honest! For a moment I thought, hey, maybe this person's numbers are lying to me... but it turned out they were not so thank you!