I already shared this in previous editions: web agents are more and more popular. They can carry out actions on the web in a window, fully autonomously.
An example:
In terms of interface, there is a conversational chat on one side and the executions on the other.
There are quite a few limits to this approach. Because the executions run in the cloud, the agent often lacks context. It is up to us to share all the information it will have available for its actions.
This is where web agents built into a browser make sense. And that is Perplexity's bet with the launch of Comet: rethinking the browsing experience in the age of AI.
We are used to using our browser to click on blue links and end up with dozens of open tabs. We get lost and often want a summary of some content, to extract information, to compare information across pages, and so on.
Our browser looks like dozens of open windows.

Comet's arrival lets you pull tabs into a conversational chat, and therefore share context to get an answer.

You can get a summary of a page in 1 click, or discuss the content of a YouTube video to find out what it covers.
Where I notice a big difference is in the integration with our tools. When I install Comet, I can import all the information from my Chrome profile (cookies, passwords, etc.). This means being already logged in to every tool: Gmail, Notion, Amazon, and so on.

No more need to share extra information (including passwords) with the models: it opens a window and performs actions while already logged in to my tools.

For example, it lets you add a conversational layer to every tool I use. I can ask which newsletters I receive by email, then unsubscribe from some of them. Or move an invitation and send an email (while keeping a human validation for sending the email).
I made a YouTube video where I share 5 concrete examples of using Comet and what really changes compared to classic web agents.
Alternatives to Comet as a browser
Other alternatives exist, such as Dia Browser, which I find more limited, because you cannot carry out actions; it is more about adding AI features that improve the experience.
OpenAI releases ChatGPT Agent to keep up
OpenAI had released, more than 6 months ago, a web agent able to carry out autonomous actions, but available only from $200/month.
Today, they go further with ChatGPT Agent, which is a combination of advanced research + operator.
The solution is a combination of several tools OpenAI already has, but it is frankly not revolutionary compared to what is already on the market for web agents or browsers like Comet.
The big plus is OpenAI expanding its features into an ecosystem of increasingly powerful tools. But nothing new under the hood based on a quick analysis.
I get the impression the tool also has a slightly larger context window, letting it carry out more complex or longer tasks, such as the ability to create files, store them on cloud services (Google Drive for example), create slides, or fill in a Google Sheets after having done a web search beforehand.
An example here on slide generation (still in beta):



