Quick answer: The same question can produce different wording or conclusions across runs.
Why it matters
Generation settings, model versions, available context and connected tools can change the response. An earlier conversation may also contain information that a fresh chat does not.
What to do next
Record your prompt, source material and the model or mode when comparing results. Use the same evaluation checklist rather than judging by style alone.