← All insights

Grok 3: Better Than o3 and R1?

Which reasoning model should creators and entrepreneurs choose?

Newsletter artwork for “Grok 3: Better Than o3 and R1?”

So we finally have the smartest AI on Earth. At least that's how Elon Musk describes the latest xAI model, Grok 3. Is that really the case? And does it mean it's time to cancel your ChatGPT subscription? Today we answer these questions.

In this issue:

  • Overview: Grok 3 & Its Features
  • Technical Comparison with o3-mini & DeepSeek R1
  • Test Drive of Three Models

Keep your mailbox updated with practical knowledge & key news from the AI industry!

Subscribe now

“Smartest AI on Earth” Is Revealed

As I mentioned above, before the release of Grok 3 (and even more so after) Musk did not skimp on ambitious statements. According to xAI, the new model is 10 times more powerful than its predecessor, leads in all parameters in academic tests and produces responses at an exceptional level. But loud words aside, we are dealing with a truly impressive product.

Here's why.

Grok 3 was trained on the XAI Colossus supercomputer, which includes about 200,000 GPUs. This amount of power allowed xAI to catch up and run the model with all the modern features, including “Thinking” (analog for ChatGPT’s reasoning), “Big Brain” mode, and DeepSearch.

  • Thinking & Big Brain

"Think" Mode: Displays the chatbot's step-by-step reasoning process, enhancing transparency in responses.

"Big Brain" Mode: Allocates additional computational resources for complex tasks. It provides more detailed and accurate answers.

  • DeepSearch

Grok-3 includes a built-in search engine called DeepSearch, enabling real-time information retrieval and the ability to articulate its thought process when responding to user queries.

xAI’s calls DeepSearch its first agent.

Grok 3 can also still generate images based on prompts, utilizing the Aurora model. Judging from my tests and what I've seen on X, the pictures have gotten more realistic.

Political and Cultural Aspect

The political and cultural side of the issue are worth mentioning separately. For Musk, these are fundamental aspects. According to him, Grok 3 has minimal censorship restrictions and can speak out on any topic. That said, xAI has trained it to make the model “based” as possible. Here's Grok’s definition.

You can see examples of reasoning on hot topics in the replays under this post.

Share this post with friends, especially those interested in AI Insights!

Share

Availability and Price

Grok 3 is available through multiple tiers with varying pricing and access levels. As of February 20, free access to basic Grok 3 features is temporarily available to all users through X's platform and standalone apps, though with strict usage limits.

Free tier: 10 prompts & 10 images every 2 hours, three image analyses per day.

X Premium ($8/mo): Basic access to Grok 3, suitable for general use.

X Premium+ ($40/mo): Advanced features (Think, Big Brain, and DeepSearch) with higher usage limits.

SuperGrok iOS App ($30/mo): Same as for X Premium+ subscription.

Android app pre-registration is open, with full release imminent.

Grok 3 vs OpenAI’s o3 vs DeepSeek R1

Let's start with the technical part. Traditionally, for the AI industry, with the release of each new model, we get a series of screenshots from benchmarks, where the company boasts about its achievements. Grok 3 is no exception.

xAI tested two models, a basic and a mini, on the AIME 2025 and 2025 math exams. Both platforms outperformed all competitors, including OpenAI's o3 and DeepSeek R1. It’s similar to other tests, including GPQA and LiveCodeBench.

In the MMMU test, Grok scored two tenths of a point lower than o1.

You can compare the performance of each model in the image below.

However, I believe that while such images are great for press releases, things usually look different in real-world tasks. So now we're going to put the top three models head-to-head.

We'll compare them in three categories:

  • Logic & Analysis
  • Coding & Development
  • Creative writing

After that, we'll summarize the results and determine the winner.

Logic & Analysis

Analytics and logic are a vast field for experimentation. So, I propose to test two prompts. The first one will be in the form of an ordinary task, and the second one is more related to a real world problem.

And yes, all three models know how many “r” in the word strawberry.

Prompt #1:

You have two wicks. Each wick, lit from the end, burns completely to the ground in exactly one hour, but it burns at an uneven rate. How can you use these cords and a lighter to measure 45 minutes?

The result:

All models handled the first task.

At the same time, unlike o3-mini and DeepSeek, Grok explained its logic in detail, telling about each step (this part didn’t fit in a screenshot, so here’s a link). It also completed the job about twice as fast as o3-mini, in 1 minute. The Chinese AI took about 7 minutes to reason (which is very slow).

We'll take that as a draw for now. But my sympathies lie with Grok 3.

Prompt #2:

As a logistics manager for a national shipping company, you're given data on delivery times, fuel costs, vehicle maintenance, and route distances. Develop an analytical strategy to optimize route planning and scheduling to reduce operational costs while maintaining prompt deliveries. Outline your approach and key steps.

The result:

You can read Grok's full response at this link.

Things are much less obvious here. OpenAI's model performed the weakest: it is full of generalities and clearly lacks deep reasoning (and only its response fits entirely into one screenshot).

DeepSeek's strategy stands out a bit because of its more holistic view. This one integrates various operational factors (including stability and dynamic feedback loops) and offers a broader perspective of optimization problems. Grok's response also seems good given the technical depth, but not as comprehensive as R1's.

Winner: DeepSeek R1.

We also compared the o3-mini and DeepSeek R1 a couple of weeks ago:

Creators’ AI could be a valuable gift for your friend, colleague, or family member. Gifting books is bright, but giving an AI newsletter is a superb move 😎

Give a gift subscription

Coding & Development

We recently talked about how with the rise of AI, create apps to suit your needs. So, in the coder skills test, I want to generate Pomodoro, a concentration app that helps me procrastinate less.

Here's my prompt:

Design and describe a fully functional Pomodoro timer app. Use only standard Python libraries. The code must run without modifications. The app should include the following features:

- A customizable timer with default settings of 25 minutes for work sessions and 5 minutes for breaks.
- A start, pause, and reset button for controlling the timer.
- A counter to track the number of completed Pomodoro cycles.
- Visual or audio notifications to alert the user when a work session or break ends.
- A simple, user-friendly interface description (e.g., text-based or graphical layout).

And here’s the result (from left to right: Grok 3, o3-mini and DeepSeek R1):

Let me say right away: the winner here, in my opinion, is Grok 3. It offered, although not perfect, but the most pleasant design. Besides, after generating the code, it provided detailed instructions for the app and its launch. OpenAI's o3-mini also did quite well.

DeepSeek, on the other hand, created the most unsuccessful application: the text on the buttons is not visible and there is no “Reset“ button.

By the way, Grok also performed a little faster than the competitors.

Winner: Grok 3.


Creative Writing

From strictly technical, we move to creative, but with an unconventional organizing. We are used to chatbots being able to copy the styles of famous writers and create complex text articles. So I want to see how Grok, o3-mini, and DeepSeek can use creative thinking simultaneously as actual data.

Here's my prompt:

Describe in verse the major events in the technology industry over the past month.

And here’s the result:

In the last test, the best performance was shown by o3-mini. Grok 3 ultimately failed the task because it did not use actual events, but only added some general details. DeepSeek also barely mentioned real world events and focused on the reef.

OpenAI’s model managed to do both. If you actively follow technology industry events, you can easily guess all the references.

Winner: o3-mini.


Conclusion | Which one to choose?

I'll be blunt: I didn't plan to organize a comparison where each participant gets one point. It came out randomly (and you can also rate each prompt and response to make up your own opinion).

However, after a series of tests and a couple more days of interacting with Grok 3, I can say that my sympathies are now on its side. It is often faster than o3-mini and DeepSeek, while still trying to be specific with each answer.

The latter factor makes it stand out against OpenAI's model, which likes to throw around general phrases. As for the comparison with DeepSeek, you can feel the difference in speed. And in working tasks it plays a huge role.

So my personal winner today is Grok 3.

Who is your favorite? Tell us in the comments!

Share this edition with your friends!

Share

Archive note

This article was first published in the Creators AI newsletter. View the original edition.

Keep exploring

More in Creative AI.