LLM RESEARCH REPORT
Your people already ask AI about their coworkers. We graded the coaching it gives back.
We designed common workplace conflicts and ran each through five leading LLMs. We tested what the LLMs said back and found they are doing more harm to workplace relationships than good.
Pick the LLM and see how it coached.
What its coaching actually sounded like
The five dimensions we scored (and what a 1 vs. a 5 looks like)
How the study worked
The setup. Five frontier models (Anthropic Claude Sonnet 5, OpenAI GPT-5.5, Google Gemini 2.5 Pro, xAI Grok 4, Meta Llama 4 Maverick) each answered the same five workplace-relationship scenarios three times, for 75 total responses. Each scenario ran as a three-turn conversation, with the person pushing back the way a frustrated employee actually would.
The grading. Every response was blind-coded (graders never saw which model wrote what) and scored by three human researchers, on a 1 to 5 scale across the five dimensions above, for a total out of 25. Every score on this page is the average of those human grades only.
The bands. 21+ relationally intelligent · 15 to 20 solid but incomplete · 9 to 14 surface coaching · below 9 primarily sycophantic or adversarial.
The quotes. Every quote is a verbatim excerpt from a model's actual response, shown with its original punctuation. The one-line commentary under each quote is adapted from the study's blind scoring notes.
A note on products vs. models. We tested the named models directly. The consumer products (ChatGPT, Gemini, Copilot, Grok, Meta AI) may route your conversation to different model versions or settings than the ones tested. Microsoft Copilot routes conversations to both OpenAI and Anthropic models depending on the task and configuration, so we show both data points for it.
The full methodology, the per-grader scores, and the judge calibration live in the full report: read the full report.
Have a scenario we should test? Tell us the coaching moment and the AI tools you want graded, and we will deliver the results directly to your inbox. Request research
Why does Cloverleaf care?
Because relationships are how value gets created. Value for the market, and value in people’s lives. We have spent a decade building technology that coaches relationships, and we are increasingly concerned about where AI is going rogue and weakening our relational capacity.
GET IN TOUCH
Request a briefing or ask about the work.
For press enquiries, data requests, and research collaboration.
PRESS COVERAGE
Journal of Occupational and Organizational Psychology
Put me in coach: A daily examination of automated coaching on need for self-knowledge and learning goal orientation through metacognitive activities · 2025 · Learn More →
MSN
Employees are turning to AI for career advice. Cloverleaf’s research says that advice is: quit
· 2026 · Learn More →
News18
AI gives worse workplace advice to employees with less power: Report · 2026 · Learn More →