【Industry News】OpenAI Releases ChatGPT Agent in the Dead of Night: Capable of Active Thinking and Tool Selection, Revolutionizing the Intelligent Agent Landscape
【Industry News】OpenAI Releases ChatGPT Agent in the Dead of Night: Capable of Active Thinking and Tool Selection, Revolutionizing the Intelligent Agent Landscape
At midnight on July 18, OpenAI held a technical live stream to unveil its groundbreaking product, Chat GPT Agent.
ChatGPT Agent possesses the ability to think and act independently, actively selecting appropriate tools from its skill repository—including Operator, Deep Research, and ChatGPT—to complete a wide range of highly complex tasks.
For example, users can ask ChatGPT Agent to analyze three competitors and create a slide presentation. ChatGPT will intelligently browse websites, select dates, filter results, run code, and even automatically generate polished slide presentations or spreadsheets.
In other words, you only need to provide a prompt, and ChatGPT Agent will handle the rest—all you have to do is wait for the results.
Full technical live stream
OpenAI CEO Sam Altman also published a lengthy article explaining ChatGPT Agent:
Today, we are launching a new product called ChatGPT Agent.
Agent represents a new height in AI system capabilities, enabling it to perform complex tasks on your behalf using its own computing power. It combines the core strengths of Deep Research and Operator, but its actual capabilities are far more powerful than they sound. It can engage in long-term thinking, use multiple tools, further refine its approach, take action, and then reflect on its actions, repeating this cycle.
For example, during the launch event, we demonstrated a presentation for preparing a friend's wedding: purchasing attire, booking travel, selecting gifts, and more.We also showed a work scenario: analyzing data and creating a presentation.
Although it is very practical, the potential risks cannot be ignored.
We have built in a large number of security safeguards and warning mechanisms, and deployed more comprehensive risk mitigation measures than ever before, covering everything from enhanced training and system protection to user control, but we cannot predict all situations.In line with our iterative deployment principle, we will issue important warnings to users while giving them the freedom to choose whether to use the features with caution.
If I were to explain this product to my family, I would say that it is at the forefront of technology and still in the experimental stage; it is an opportunity to experience the future, but we do not recommend using it for high-risk tasks or scenarios involving large amounts of personal information until we have researched and improved it through practical application.
We do not know exactly what impacts it may have, but malicious actors may attempt to “trick” users' AI agents into disclosing private information that should not be disclosed or performing actions that should not be performed, and these methods are beyond our current ability to predict. We recommend that, to minimize privacy and security risks, users only grant the agent the minimum permissions necessary to complete tasks.
For example, I can allow the Agent to access my calendar to find a suitable time for a group dinner. However, if I only want it to help me buy some clothes, no additional permissions are needed.
Tasks like reviewing the emails I received last night and autonomously handling all the necessary actions without further prompting carry higher risks. This could lead to untrusted content in malicious emails deceiving the model, resulting in data leakage.
We believe that learning from real-world applications is critical, and that people should adopt these tools cautiously and incrementally as we better quantify and mitigate potential risks. As with other new capability tiers, societal, technical, and risk mitigation strategies need to evolve in tandem.
Technically, ChatGPT Agents process tasks through their virtual computers, enabling them to switch seamlessly between reasoning and execution.When faced with complex tasks, it can not only perform logical reasoning but also execute tasks, thereby independently completing complex multi-step tasks.
For example, when a user asks ChatGPT Agent to “check my calendar and provide a brief report on upcoming customer meetings based on the latest updates,” it can understand the task requirements, proactively retrieve information from the calendar application, and organize concise report content.
Another important feature of ChatGPT Agent is its multi-tool integration capability, which combines the website interaction capabilities of Operator, the information integration capabilities of Deep Research, and the deep conversation capabilities of ChatGPT to form a unified intelligent agent system.
Operator enables ChatGPT intelligent agents to scroll, click, and enter text on web pages, allowing them to interact directly with websites.while Deep Research excels at analyzing and summarizing information, helping ChatGPT intelligent agents handle complex multi-step tasks.
Additionally, ChatGPT Agent is equipped with various network tools, including a visual browser, text browser, and direct API access permissions. These tools provide ChatGPT intelligent agents with different pathways for accessing and interacting with network information, enabling them to select the optimal path to complete tasks efficiently.
For example, financial data or sports scores can be quickly retrieved via API, while visual interaction with web pages designed primarily for humans is also possible. All these operations are performed within ChatGPT's own computational environment, ensuring that all relevant background information is shared throughout the task, regardless of the tool combination used.
During task execution, ChatGPT agents can dynamically learn and optimize their working methods. Through reinforcement learning, the model adjusts its strategies based on task outcomes, continuously improving its performance.This dynamic learning ability enables ChatGPT agents to flexibly adjust their action strategies according to different task requirements, improving task completion speed and accuracy. ChatGPT Agent is also designed for iterative and collaborative workflows, significantly improving its interactivity and flexibility. During task execution, users can interrupt the conversation at any time to clarify instructions, reorient the task, or guide it toward the desired result.
ChatGPT AI will resume from where it left off, integrating new information without losing prior progress. This allows users to adjust task direction at any time during execution, ensuring that task outcomes align with user expectations.
In terms of security, ChatGPT AI is designed with user safety in mind. Before performing sensitive or critical operations, ChatGPT explicitly seeks user authorization to ensure users maintain control at all times.Additionally, ChatGPT AI agents feature proactive monitoring and risk mitigation capabilities, enabling them to actively reject high-risk tasks, such as financial transactions or sensitive legal interactions.
According to test data published by OpenAI, ChatGPT Agent demonstrated outstanding performance across multiple benchmarks. In the “Human Ultimate Exam,” it achieved a pass rate of 41.6% in a single attempt, setting a new state-of-the-art (SOTA) record, which improved to 44.4% when using a parallel strategy;In the “Frontier Mathematics” benchmark, it achieved an accuracy rate of 27.4%, significantly outperforming previous models.
In internal benchmark tests simulating complex real-world tasks, the output was comparable to or better than humans in approximately half of the cases for complex and economically valuable knowledge-intensive tasks, significantly outperforming o3 and o4-mini, covering a wide range of real-world professional tasks.
In DSBench, it significantly outperformed humans; in SpreadsheetBench, it significantly outperformed existing models, achieving a score of 45.5% when given the ability to directly edit spreadsheets, far exceeding Copilot's 20.0% in Excel.
In an internal benchmark measuring the modeling task capabilities of investment bank analysts, it significantly outperformed deep research and o3, covering multiple modeling tasks, all scored against hundreds of criteria.
In the BrowseComp benchmark, it achieved a SOTA score of 68.9%, 17.4% higher than Deep Research; in WebArena, it outperformed CUA driven by o3.
Some users have commented that ChatGPT Agent feels more like Manus 2.0. While Manus was indeed an interesting concept when it was first launched, it was too unstable to be used effectively.
I’m really looking forward to trying ChatGPT Agent to see if it lives up to the hype. Is this another step toward AGI?
This is truly exciting, and I can’t wait to try it. I fully agree with this approach: “Powerful agents may possess extraordinary capabilities, but they also come with significant risks. These risks stem not only from malicious attackers but also from hallucination issues. Let’s explore together and gain a deeper understanding of their underlying impacts.”
The team's update is fantastic, and I'm very excited about it. I can't wait to use it and see it grow stronger over time.
I really appreciate that you're putting this in our hands now, rather than waiting for some unattainable zero-risk standard. In my view, trusting users with reminders and precautions is the right approach.
This is incredible! Seeing AI actually browse websites and complete real-world tasks feels like science fiction coming to life. I'm already thinking about how this could simplify workflows for content creators and small businesses. The productivity revolution starts now!
Comments
Post a Comment