← All insights

OpenAI o3 vs DeepSeek R1

OpenAI is fighting back, but which model is better?

Newsletter artwork for “OpenAI o3 vs DeepSeek R1”

At the end of last week, OpenAI finally opened access to its o3-mini model (for all users!), which responded to the buzz from DeepSeek. Today, we review it, compare it to the DeepSeek R1, and determine which reasoning model is right for your tasks.

We'll run a series of tests in which each participant solves complex problems from different domains. Let's go!

Keep your mailbox updated with practical knowledge & key news from the AI industry!

Subscribe now

OpenAI o3-mini | Quick Overview

Although we are comparing o3 with DeepSeek R1 here, we won’t be taking a detailed look at the Chinese startup in this post. If you want to learn more about it:

OpenAI's o3-mini is the latest addition to the reasoning model lineup, building upon the foundation laid by the o1 model. While o1 was designed to handle complex tasks across various domains, o3-mini introduces enhanced reasoning capabilities, particularly in science, technology, engineering, and mathematics (STEM) fields.

In performance evaluations, o3-mini has improved over its predecessor. It matches o1's proficiency in tasks like mathematics and coding reasoning but delivers responses approximately 24% faster. Expert assessments indicate that o3-mini provides more accurate answers, with a 39% reduction in major errors on challenging questions.

As for availability (I guess we can thank DeepSeek for this), o3-mini is available to all categories of ChatGPT users, including the free tier. However, depending on the type of subscription, OpenAI will give you a different number of messages per day.

Pricing

  • ChatGPT Pro ($200/month): Unlimited use
  • Plus ($20/month): 150 daily messages
  • Free users: Limited testing via "Reason" mode (but less than 150)
  • API Cost: $1.10 per million input tokens and $4.40 per million output tokens
Overall, the cost for APIs in o3-mini is 63% cheaper than o1-mini.

As we gradually move to comparison, DeepSeek's R1 presents a notably more economical option. DeepSeek R1 is priced at $0.14 per million input tokens (cache hit), $0.55 per million input tokens (cache miss), and $2.19 per million output tokens.


o3-mini vs DeepSeek R1

Now, about the technical side of the question. We won't go into all the benchmarks and analyze dozens of synthetic tests. While they give a general idea of performance, these numbers are useless for real-world scenarios.

Here are the LiveBench data:

LiveBench updates its questions using recent sources such as arXiv papers, news articles, and IMDb movie synopses. Each question has an objective, verifiable answer, allowing automatic scoring without relying on human or LLM judges.

As you can see, the OpenAI model slightly outperforms its competitor in almost all checks. The only exception is math reasoning (76.55 vs. 79.54), where DeepSeek R1 is better. At the same time, it is important to keep in mind that we are talking about o3-mini (high)—the most productive subcategory.

Approximately the same results show other benchmarks:

Very close, but the o3-mini wins more often: in four tests out of six.

API

To avoid burdening you with unnecessary details, below is a table comparing the APIs of the two models. Considering the cost, DeepSeek could be more attractive, but we shouldn’t forget about the problems with data leaks and the fact that reliability leaves much to be desired (unlike o3-mini).

Who stole from who?

There is also one funny thing. After the release of DeepSeek, some users on X noticed that the model sometimes called itself “ChatGPT.” This led many to suspect that the Chinese AI is based on competitor data. But that's not all.

o3-mini distinguished itself, occasionally reasoning in Chinese. So, now many users think that OpenAI copied some of DeepSeek's open-source code to release an updated model as soon as possible but didn't have time to edit the code.

This information cannot be reliably confirmed, so you decide which is true.

Prompts Testing | Who’s better?

We will compare the o3-mini and DeepSeek R1 within five complex and practical prompts, defining the best AI model on the market. Here are our categories:

  • Creative Writing
  • Analysis of Real-Time Content
  • Logical Problem Solving
  • Financial Modeling & Projections
  • Business Strategy

Within these topics, the emphasis will be on use cases for creators and entrepreneurs. At the end of each test, we’ll look at the results and determine the winner.


1. Creative Writing

Prompt:

You are a founder of a tech startup. Write a first-person investor memo convincing venture capitalists to invest in your human-AI collaboration platform. The memo should include:

- A compelling opening hook
- A detailed problem statement with real-world consequences
- How your AI solution works (be specific)
- Competitive landscape analysis with major risks
- Closing argument with a clear CTA for investors

First test, but we can already see an interesting difference in approach. While o3-mini decided to inundate investors with a set of rather general (and, let's be honest, watery) theses, DeepSeek R1 pitched the project with numbers, data and enterprises.

This is clearly evident even from the first paragraph. Instead of a compelling and realistic hook, OpenAI starts by communicating a breakthrough but doesn't specify what it will be. The Chinese model, in my opinion, does a better job: it starts with a significant problem in the labor market, backed up by data from a reputable firm, and concludes with an ambitious goal.

I also liked that DeepSeek came up with a name for the startup.

Winner: DeepSeek


2. Analysis of Real-Time Content

Prompt:

Find the three most-discussed AI news in the past week based on TechCrunch, The Verge, and X (Twitter). Summarize them in a structured way, including:

- The basic concept behind each trend
- Why these news attracted the attention 
- What market events might follow them

Fit it into three short paragraphs
Focus: Ability to extract insights from real-time data and summarize effectively.

Now, let's see how the reasoning models perform with AI search.

Here, the results are less unambiguous, and we can see the difference in the two models' approaches. o3-mini took the query literally, so it used one source for each news story: one from TechCrunch, one from The Verge, and one from X.

This made the first two choices—new AI policies from Meta and the emergence of DeepSeek—good, but the third was some obscure insight from a not-so-popular account. With a system like that, it will be hard to stay on trend.

At the same time, R1 chose the release of OpenAI's Operator, Microsoft's restructuring, and the Stargate Project discussion as topics. However, the Chinese model didn’t use my suggested sources and turned to third-party sites. And, unlike o3-mini, DeepSeek R1 forgot about the main topic of the week: its own release.

Apparently, the Chinese model is shy, so it downplayed its own merits.

Overall, I found DeepSeek's answer more convincing.

Winner: DeepSeek.

Creators’ AI could be a valuable gift for your friend, colleague, or family member. Gifting books is bright, but giving an AI newsletter is a superb move 😎

Give a gift subscription

3. Logical Problem Solving

Prompt:

A founder is running an AI-powered legal document automation startup. They are struggling with:

– High customer acquisition costs

– Low customer retention after 3 months

– Competition from open-source alternatives

Suggest a step-by-step structured turnaround strategy using first-principles thinking. Prioritize based on impact vs. execution difficulty. And make your answer as short as possible.
Focus: Logical reasoning, business acumen, and structured thinking.

I think the o3-mini is starting to get better.

While DeepSeek’s approach offers commendable elements—especially the targeted retention fixes and the idea of a freemium tier—the o3-mini model’s strategy builds on a framework that addresses the entire customer lifecycle.

This plan could create a more sustainable competitive edge by prioritizing value clarification and leveraging a combination of inbound marketing, strong partnerships, and differentiated product features.

DeepSeek's ideas add a helpful perspective, particularly in customer retention tactics. However, in my opinion, OpenAI's model gives comprehensive approach that provides a more strategic pathway.

Winner: o3-mini.


4. Product-Market Fit Validation

The answers turned out to be pretty big.
The full versions of both models are on Google Docs.

Prompt:

You are launching a new AI-powered productivity tool for solopreneurs.

Outline a 30-day go-to-market plan, including:

Fastest way to validate product-market fit
Customer acquisition strategy (with estimated cost per acquisition)
Key engagement and retention metrics
Early monetization options (freemium, SaaS, etc.)
Focus: GTM strategy, market validation, growth hacking.

It's very close here, but I'm leaning toward the o3-mini again. The OpenAI model offers a more flexible methodology that focuses on rapid feedback and continuous iteration. This framework allows for faster customization based on real-time user data, a significant advantage in the AI market.

The o3-mini model's emphasis on iterative development and rapid time-to-market is desirable to solopreneurs because it ensures that the product can evolve quickly to meet changing needs.

Winner: o3-mini.

By the way, we have a great toolkit for solopreneurs:

5. Business Strategy

Full versions of the responses are here.

Prompt:

Assume you are an AI-driven GPT wrapper startup competing in 2025.

- Find and analyze the three biggest GPT wrapper companies right now.
- Identify their unique selling propositions (USP).
- Suggest a defensible moat for a new entrant.
- Model a customer acquisition strategy with estimated CAC and LTV.
Focus: Market analysis, strategic thinking, competition mapping.

It's time to decide on the ultimate winner.

Both the DeepSeek answer and the o3-mini performed well in the final round. Both models walked through key areas-analyzing leading companies, identifying unique sales propositions, describing a defensive moat, and developing a customer engagement strategy-but they differ in emphasis and depth.

The OpenAI model not only analyzed leading competitors but also proposed a strategy, targeting underserved niches and creating proprietary integrations. In addition, the customer engagement strategy was laid out with clearer metrics and phased execution.

This made the o3-mini answer more appealing to me. But you can check out the full answers and form your own opinion.

Winner: o3-mini.

If you want to learn more about AI wrappers, check out this post:

Final Thoughts

The result: o3-mini beats DeepSeek in our testing (3 vs. 2).

Both models performed well, and the gap between them isn’t huge. This wasn’t a perfectly objective test (nor did I try to make it one), and I wouldn’t be surprised if a different set of questions flipped the result in favor of DeepSeek R1.

What really stands out is the progress we’re seeing in AI reasoning. Just a year ago, any of the tasks above would have required a dozen prompts (and some would not have been solved at all) and edits, but now we get those answers within a couple of minutes. And that is nothing short of gratifying.

That said, they’re still far from perfect. Both o3-mini and DeepSeek R1 had moments where their reasoning wasn’t quite there yet. There’s a lot of room for improvement, which makes this space exciting to watch.

For now, we wait. With future updates, the balance may shift again. Either way, I’ll be keeping an eye on both—if there’s one sure thing, AI isn’t done evolving yet.

Share this edition with your friends!

Share

Archive note

This article was first published in the Creators AI newsletter. View the original edition.

Keep exploring

More in Creative AI.