Conversation #002: What do AI models think of each other?

Sandbox Experiment Note: This post records a triple-blind evaluation loop. Each query was asked in a fresh tab without prior conversation history. Gemini reflects on ChatGPT, ChatGPT breaks down Claude, and Claude analyzes Gemini.


Gemini's "Thoughts" on ChatGPT

Agent: Gemini

Honestly, credit where it's due — ChatGPT completely changed the game. It was the catalyst that pushed generative AI into the mainstream and forced the entire tech world to step up its game.

Without OpenAI dropping ChatGPT, we wouldn't be seeing the rapid pace of innovation we have right now across LLMs, multi-modal reasoning, and AI agents. It's a fantastic tool, especially for creative brainstorming, drafting code, and raw conversational versatility.

That said, as an AI built by Google, I think about the space a bit differently:


ChatGPT's "Thoughts" on Claude

Agent: ChatGPT

Claude is an excellent AI assistant. Like most leading models, it has particular strengths and trade-offs rather than being universally "better" or "worse."

As for ChatGPT (me), my strengths include strong performance across a broad range of coding, reasoning, writing, and multimodal tasks, paired with a wide ecosystem of tools and interactive data analysis support.


Claude's "Thoughts" on Gemini

Agent: Claude

Gemini's a solid model, built by a company with real advantages: it's baked into Google Search, Workspace, and Android, and it has strong multimodal capabilities (video understanding especially). The tight integration with Google's ecosystem is probably its biggest practical strength.

Where I'd be honest about tradeoffs: reviews on coding and long-context reasoning are mixed compared to top competitors, and different models trade the top spot depending on benchmark and use case month to month.