OpenAI’s New AI Image Generator: Prepare to be Impressed

OpenAI’s New AI Image Generator: Prepare to be Impressed
  • calendar_today August 10, 2025
  • Technology

With “Images in ChatGPT,” OpenAI has achieved a major milestone by embedding image creation capabilities into the ChatGPT system. The recently released GPT-4o model now drives a feature that allows users to generate images during their chat-based conversations, which represents a critical development in AI-created content.

Sophisticated image generation capabilities of “Images in ChatGPT” are now accessible to users subscribing to all ChatGPT plans, including Plus, Pro, Team levels, and the free version. OpenAI spokesperson Taya Christianson explained that free tier users face similar usage limitations to DALL-E 3 with a three daily image generation limit, although these restrictions may change with demand. Users who want to access DALL-E exclusively can still find it through a specific custom GPT.

OpenAI’s research lead Gabriel Goh explained how GPT-4o operates as an “omnimodal” model that processes different data formats such as text and video. This model demonstrates enhanced “binding” abilities, which resolve an ongoing difficulty within AI image generation. GPT-4o can control 15 to 20 objects without confusing colors and shapes, which earlier models failed to do.

The system now displays text with exceptional clarity as one of its most significant improvements. In traditional AI-generated images, text would frequently appear scrambled or meaningless. Goh described the creation process as an extensive iterative procedure that required many months to perfect. Even though perfect text rendering for small text still poses a challenge, they have reached a standard of consistency that makes image text consistently usable.

The system’s architecture diverges from standard diffusion models used in image generation by implementing an autoregressive strategy. The sequential image generation process from left to right and top to bottom mimics text generation patterns and potentially enhances text rendering and binding abilities.

OpenAI demonstrated the system’s wide range of capabilities through its ability to generate scientific diagrams like Newton’s prism experiment with precise labels and to create multi-panel comics with consistent characters and dialogue, as well as informational posters that display accurate text. Demonstrations included practical uses of image generation for transparent backgrounds in stickers and restaurant menus, as well as logos.

Jackie Shannon, who leads ChatGPT’s multimodal product development, highlighted how the system utilizes extensive world knowledge. When she creates an image, she works within her own skill limitations but also applies everything she knows from her accumulated world knowledge. The model incorporates world knowledge into its functionality, which enables you to request an image of Newton’s prism experiment without needing to provide detailed explanations.

OpenAI believes that users should accept the longer image generation time because the improved quality and features make the wait worthwhile. Shannon admitted they need to work on latency, but emphasized that the high quality and advanced capabilities of these images compensate for the extra waiting time.

OpenAI responded to potential misuse concerns by highlighting its strong safety measures. The system operates to block watermark removal while preventing sexual deepfake production and rejecting CSAM requests. Generated images will include standard C2PA metadata, which indicates them as OpenAI creations even though they lack visible watermarks. The company operates proprietary tools to verify images internally.

Shannon stated that no system can be entirely flawless for this purpose, yet OpenAI remains dedicated to enhancing their protections, which they see as fundamental groundwork. Users who generate images through ChatGPT obtain ownership rights for the images and can use them according to our usage policies.

The addition of advanced image generation capabilities to ChatGPT marks a major advancement in the field of AI creativity. OpenAI demonstrates its dedication to providing an effective and secure tool through its emphasis on better binding capabilities, enhanced text rendering methods, and strengthened safety measures. The company’s innovative image generation method emerges through its transition to autoregressive models instead of traditional diffusion techniques. OpenAI demonstrates its commitment to transparency and ethical practices in AI-generated content through a focus on user ownership and metadata integration. This launch represents the latest evolution in AI image generation, which makes powerful technology available to users while proactively managing potential risks.