OpenAI o1 Review: It's Not So Perfect?
Discussing o1 from the perspective of creators and entrepreneurs

Hi!
So, last week, OpenAI finally let some users try their new model. Traditionally, some users were delighted (or even called o1-preview “that GPT-5”), while others were disappointed by significant limitations.
Today, let's discuss the new OpenAI model, its advantages and disadvantages, and run relevant tests for creators and entrepreneurs. Ultimately, we'll try to answer whether OpenAI succeeded in creating the perfect model.
In this edition:
- Opportunities and limitations of o1-preview
- The technical performance of the new OpenAI model
- Testing o1-preview from a founder's and creator's perspective
- Tips & Tricks for working with o1
Great Opportunities & Limitations

What is OpenAI's o1-preview?
It is a new model developed under the code name “Strawberry.” (Technically, it is a family of models because there is also o1-mini, but in our case, it is not so important.) Its distinguishing feature is a tendency to reason and, in a sense, self-reflect. It takes more time to think and give the answers.
This makes o1-preview noticeably slower than GPT-4 but allows it to reason about complex tasks and solve complex problems better than its predecessor. Therefore, it is primarily focused on solving not primitive queries but working in science, coding, math, and other complex domains.
Here is how its developers evaluate o1:
In our tests, the next model update performs similarly to PhD students on challenging benchmark tasks in physics, chemistry, and biology. We also found that it excels in math and coding. In a qualifying exam for the International Mathematics Olympiad (IMO), GPT-4o correctly solved only 13% of problems, while the reasoning model scored 83%.
According to OpenAI's research lead, Jerry Tworek, o1 has been trained using a completely new optimization algorithm and a new training dataset tailored explicitly for it. The model processes requests using a “chain of thought,” similar to how people process problems by going through them step by step.
OpenAI also says that this o1 is a new level of AI capability from which the company now counts its progress (that's why it's called “o1”, by the way).
Keep your mailbox updated with practical knowledge & key news from the AI industry!

According to tests conducted by OpenAI, o1 ranks in the 89th percentile on Competitive Programming (Codeforces), ranks among the top 500 students in the U.S. Mathematics Olympiad (AIME) qualifying round, and surpasses human Ph.D. level accuracy on the Physics, Biology, and Chemistry (GPQA) problem test.
The o1 also significantly outperforms the GPT-4o on various benchmarks, including the 54/57 MMLU subcategory. The image above shows that we are talking about a multiple advantage.
Sharing is caring! Refer someone who started a learning Journey in AI!
As for the limitations, they are quite serious.
For starters, o1 does not support file handling. Unlike GPT -4, you can't provide image models or documents. Online search doesn't work either, so you have to pay special attention to the accuracy of the quick response to determine if the information is from the latest sources. There are also some limitations on the API. Developers must be at level 5 to gain access.
Another factor that makes you consider whether o1-preview is worth working with is its cost. The o1-preview API costs $15 for 1M input tokens or text snippets analyzed by the model and $60 for 1M output tokens. By comparison, its predecessor, GPT-4o, costs $5 for 1M input tokens and $15 for 1M output tokens.
After that, perhaps the most infuriating limitation is the number of messages that can be sent. The standard o1 model only supports 30 messages per week. You'll have to count the number of them yourself and plan your session well.
To be fair: there are no message limits in the API.
So that you don't have to spend too much time counting every post, we've done some tests and are sharing the results. In the processes, we had several goals: to find out whether o1 is suitable for non-scientists, how good it is in application tasks, and, in the end, whether it makes sense to overpay or you should stop at GPT-4o.
Testing OpenAI's o1-preview
First, we have to define from which positions we will test o1. OpenAI claims that, among other things, the new model helps solve problems in science, including biology, physics, and many others. Unfortunately (or fortunately), we are not scientists, so we will test o1 in more applied problems, for example, on behalf of an entrepreneur/creator with several issues to solve.
During testing, we will be comparing o1 to GPT-4o.
Our test simulates the preparation for an AI startup's entry into a new market. We already have a company that has been successful in North America and is now preparing to launch in Southeast Asia. Here's what our prompt looks like:
I intend to enter the Southeast Asian markets with my AI healthcare startup that has succeeded in North America, offering personalized AI-powered healthcare solutions. Develop a brief strategy
Here is the result we got from o1:
You can view the full conversation using this link.
And this is what the GPT-4o generated:
Here's a link to our conversation with GPT-4o.
As you can see, o1 offered a much more detailed strategy and, as a bonus, showed his reasoning. The process took about 10 seconds (in other tests, prompts required about 15-20 seconds for reasoning), which doesn't look like a problem.
Regarding content, o1 has also shown me to be more productive. For example, in strategies from both bots, we can see a paragraph about potential partners. However, while GPT-4o suggested quite generic medical clinics and insurance companies, o1 also advised approaching government organizations and technology companies.
Overall, the strategy suggested by the new model is more like a plan that can be evolved. GPT-4o, on the other hand, just gave some general advice, and if you want a better result, you will have to rewrite your prompt or provide a lot of additional data.
Moving on to the next test.
Now, let's look at a more typical business problem. We have an e-commerce brand that urgently needs a content plan for various social networks a month ahead.
Here’s our prompt:
I am launching an ecommerce business selling designer clothing from scratch. To make my brand recognizable, I need a solid content plan for LinkedIn, Instagram, and TikTok for the next month. Make one and explain your decisions.
The results of both models were very voluminous so I won't add screenshots. Instead, I will give you links to conversations with o1 and GPT-4o.
- Conversation about a plan for an eCommerce brand with o1-preview.
- Conversation about a plan for an eCommerce brand with GPT-4o.
Here, the difference is much more noticeable.
The task was more complex, so o1 took 27 seconds to answer. Slow but acceptable. In addition, it gave concrete suggestions and explained every point, including content, visuals, and purpose. Also, the plan itself is very diverse.
For example, the new model considered that LinkedIn, Instagram, and TikTok require different numbers of posts. Plus, it explained each solution in detail, like content themes and objectives, engagement tactics, and cross-promotion. GPT-4o instead said that professionals use LinkedIn and TikTok was created for fun videos. Well, thanks.
Summarizing the test results, we can conclude that o1 is much better than GPT-4o. This is partly true; reasoning gives a big advantage over the still-current models. Nevertheless, if you have a lot of interaction with ChatGPT, Claude, and others, you know that it can be difficult to get a perfect result the first time, but if given enough information, these models can also offer good things for your problems.
Tips & Tricks for Your Workflow
While testing o1, we came up with a few points to remember while working with it. They will help you find common ground with the new model and get the results you want noticeably faster (considering you'll only have 30-50 posts per week, I suggest familiarizing yourself with them).
Here's what we've identified:
- Prompt engineering is not required. This may seem strange since we've spent so much time getting used to this communication method with the LLM. However, when working with o1-preview, prompt engineering can worsen the result. In addition, there is no longer a need to use cues that trigger thought chain activity because it happens automatically.
- Limit additional context for advanced search generation (RAG). This is not only our advice but also that of OpenAI. Adding more context or documents when using models for RAG tasks can overcomplicate the response.
- The context should be provided in the task ticket format. Imagine that you are assigning a task to a colleague in Jira and proceeding similarly when working with o1.
- **Content delimiters for your context ``, \*\* and <tag>your specific text</tag> are a must:** The model should understand when you give data for an “example.”
And, of course, remember that o1 is still error-prone, just like any other SOTA-level model. So, double-check all answers and refine your queries if necessary.
Conclusion
I guess it's time to answer the question from our title. And it's pretty obvious: no, o1-preview is not perfect. It is not a product for everyone, even if you are deeply immersed in the AI industry and used to testing every model right after release.
At the same time, o1 is the best model for marketing strategies and research right now. It's always trying to produce quality results with even a few inputs. I would say that o1 has the most significant prospects to become the ideal model. But before that, OpenAI should open up access to file upload and search and increase the number of available posts.
So, the conclusion is this: if you need a good model for work tasks - choose GPT-4o. It's still relevant and universally suitable for various creators and entrepreneurs. If you want to solve incredibly complex tasks (for example, those related to math and complex sciences), try OpenAI o1.
Regarding developers and APIs, I would also take my time. The model is costly right now, which could hinder development, especially coupled with limitations on the number of messages and functionality. On the other hand, with o1, you can get the best results on the LLMs market. The choice is yours.
Share this edition with your friends!
This article was first published in the Creators AI newsletter. View the original edition.


