Can we really achieve consistent & deterministic response from AI models?
Situation when an LLM produces very similar response (but not strictly
the same) to the same prompt has long seemed to be a feature. Still,
it is annoying and makes LLM output feel unreliable. In this video we
are using a technique called "Batch invariance" to force the same
response to the same prompt.