This reads poorly to me. So they’ve proven the models CAN work, but they also say in the next line that they CANNOT do high value work in production.
No other comments from me but that first two sentence opener should have been massaged a bit. I’m sure they have a PR / messaging person (agent?) though, so maybe I’m reading too far into it.
AI support has been terrible so far in my experience and tends to over optimize for a happy path easily solved on a company website rather than solve the issue that made me call them on the phone that is likely a system bug or other issue.
Building your own enterprise chatbot with in-house domain experts is the only path that makes sense to me. The software piece is really not that difficult. There are a lot of examples and options to pull from now. You could maybe implement a custom MCP server and use the M365 copilot on top if you don't want to reinvent the wheel.
I think bringing in OAI consultants is probably a mistake in most cases. I've already seen one AI consulting team and they just cannot get deep enough fast enough. It would take them years of suffering our codebase and daily procedures to get to the point where they could actually make an impact.
I guess my question is, what is a safe AI "process percentile" for a company that lets it recover if the AI goes down or goes haywire? No one knows (?), but it's worth thinking about.
I am also curious just how much direction the AI will need given the environment is dynamic. Humans generally align to the company's goals because they are rewarded (paid) to, and need to to survive. Relationships are built on unspoken and spoken communication. Unspoken could be cultural, hierarchical, environmental pressures, etc. As an extension of that, companies build relationships with other companies by sending (essentially) diplomats and ambassadors to each other (the article mentioned using AI for sales). -- All this to say: if you have to constantly explain the job to someone, they are not fit for the job.
We'll see, I guess.
I called my window contractor this week, I need new skylights. Instead Janet answering, they now use an AI operator-lead-generator thing.
After a few of minutes of answering questions, I just hung up. I’m unsure why, I don’t miss Janet or even genuinely like talking to people, and Im definitely not anti-AI. I just intensely disliked the experience, enough to walk away from a vendor with whom I’ve done tens of $ks in business (I found another window guy).
Otherwise, I see this as potentially useful for firms that dont have global pressence and want to service a support system across multiple timezones where it might be difficult to staff a night-shift. However, issues of trust and quality will always be something to contend with here.
That’s the first I’ve heard of it. Has anyone tried it?
Is this the correct reading? If so, I'm not that impressed.
If you call a bar in SF or a concierge in Vegas, you'll often get a voice chat agent. There's no terms of service you agree to. In aggregate, it's a pretty rich dataset on openai's end.
What happens to your voice, your queries, your PII?
I don't want to laugh too hard this afternoon. But more seriously, I'm not really sure who this product is for in a way an existing tool can't handle? Like if a business really wants to go deep on agentic workflows for customer service, what is going to make them reach for this?
This has a lot of words with very little information. Having codex make changes to code because someone made a request to IT is insane, even if it needs approval.
Idk man, I feel like I am taking crazy pills.
I use codex everyday but I plan the work and codex performs the work bit by bit, which allows me to review every piece of code. I would not sleep well not knowing what is in production.
Is this some sort of joke? Like. That’s the proof. If it can do reliable work then that’s the proof. Contradiction in sentence one. If your pipeline breaks 80% of the time but sometimes the magic token lottery gives you something remotely useful, that’s NOT proof. The ceiling for these companies is so low it’s unbelievable. What happened to products that worked first, and worried about the rest later? Can we have an inspired ad article for that instead?
or again it's simply OpenAI looking for PMF to justify their valuations ?
I think it's the latter and will fail again just like Sora.
Did it use OpenAI presence for that?
I'd bet in 2 weeks nobody will remember this
Not sure if they're just trying to find what the abstractions are worth focusing, or if it will stay that way, but certainly cool to see how much they're trying out.