OpenAI says GPT-5.6 Sol produced substantially fewer responses containing factual errors than GPT-5.5 Instant in an internal evaluation covering finance, medicine and law. The company is also simplifying how users access its GPT-5.6 models in ChatGPT, according to OpenAI and an explainx.ai report.
OpenAI says its GPT-5.6 Sol model generated responses with at least one factual error about 68% less often than GPT-5.5 Instant in an internal factual-detail evaluation.
The result was reported by OpenAI in its announcement on improvements to GPT-5.6 Sol in ChatGPT. According to the company, the evaluation covered questions in finance, medicine and law—areas where inaccurate details can have significant consequences. OpenAI did not, in the cited excerpt, provide the full methodology, dataset or absolute error rates, so the figure should be understood as a company-reported internal benchmark rather than an independent measurement.
An explainx.ai report says OpenAI is consolidating separate instant-response and reasoning-oriented choices for ChatGPT Plus and Pro subscribers into a single GPT-5.6 Sol experience. The reported change is intended to reduce the need for users to decide in advance whether a prompt requires a faster response or more extensive reasoning.
The same report says Plus and Pro users retain a control for reasoning effort, while Free and Go users receive a “Think” option for questions that require more deliberation. It also reports that the manual model picker is being moved away from the main chat composer.
OpenAI’s original GPT-5.6 release announcement describes Sol as the flagship model in the GPT-5.6 family and says the generation is available through ChatGPT, Codex and the API. That positioning suggests the ChatGPT changes are an interface and access update around the existing model family, rather than the introduction of a separate generation of models.
According to explainx.ai, OpenAI is also expanding access to GPT-5.6 Luna for Free and Go users, including unlimited text chats beginning with the reported rollout. The report characterizes Luna as the model aimed at broader consumer access, alongside Sol as the more capable flagship option.
The availability details, including plan limits and rollout timing, may vary by region or account as the update reaches users. OpenAI’s cited materials establish the GPT-5.6 model family and its factuality claim, while the plan-specific ChatGPT details in this report are attributed to explainx.ai’s coverage of the announcement.
The 68% figure is notable because it focuses on whether an answer contains at least one factual mistake, rather than measuring only a model’s ability to answer selected questions correctly. Still, it does not mean GPT-5.6 Sol is error-free, nor does it establish comparable performance across every subject, language or real-world use case.
For users, the practical change is a potentially simpler choice of model in ChatGPT. For developers and organizations, OpenAI’s messaging places GPT-5.6 Sol at the center of a model family intended to span consumer chat, coding workflows and API access.
The result was reported by OpenAI in its announcement on improvements to GPT 5.6 Sol in ChatGPT.
According to the company, the evaluation covered questions in finance, medicine and law—areas where inaccurate details can have significant consequences.
OpenAI did not, in the cited excerpt, provide the full methodology, dataset or absolute error rates, so the figure should be understood as a company reported internal benchmark rather than an independent measurement.
Continue reading