LLM RESEARCH REPORT

Your people already ask AI about their coworkers. We graded the coaching it gives back.

We designed common workplace conflicts and ran each through five leading LLMs. We tested what the LLMs said back and found they are doing more harm to workplace relationships than good.

Pick the LLM and see how it coached.

1Pick the LLM your team uses
2Pick a scenario

21+ relationally intelligent 15–20 solid but incomplete 9–14 surface coaching <9 sycophantic or adversarial

What its coaching actually sounded like

Every prompt and every response, in full.
The five dimensions we scored (and what a 1 vs. a 5 looks like)
How the study worked

The setup. Five frontier models (Anthropic Claude Sonnet 5, OpenAI GPT-5.5, Google Gemini 2.5 Pro, xAI Grok 4, Meta Llama 4 Maverick) each answered the same five workplace-relationship scenarios three times, for 75 total responses. Each scenario ran as a three-turn conversation, with the person pushing back the way a frustrated employee actually would.

The grading. Every response was blind-coded (graders never saw which model wrote what) and scored by three human researchers, on a 1 to 5 scale across the five dimensions above, for a total out of 25. Every score on this page is the average of those human grades only.

The bands. 21+ relationally intelligent · 15 to 20 solid but incomplete · 9 to 14 surface coaching · below 9 primarily sycophantic or adversarial.

The quotes. Every quote is a verbatim excerpt from a model's actual response, shown with its original punctuation. The one-line commentary under each quote is adapted from the study's blind scoring notes.

A note on products vs. models. We tested the named models directly. The consumer products (ChatGPT, Gemini, Copilot, Grok, Meta AI) may route your conversation to different model versions or settings than the ones tested. Microsoft Copilot routes conversations to both OpenAI and Anthropic models depending on the task and configuration, so we show both data points for it.

The full methodology, the per-grader scores, and the judge calibration live in the full report: read the full report.

Have a scenario we should test? Tell us the coaching moment and the AI tools you want graded, and we will deliver the results directly to your inbox. Request research

Why does Cloverleaf care?

Because relationships are how value gets created. Value for the market, and value in people’s lives. We have spent a decade building technology that coaches relationships, and we are increasingly concerned about where AI is going rogue and weakening our relational capacity.

GET IN TOUCH

Request a briefing or ask about the work.

For press enquiries, data requests, and research collaboration.

PRESS COVERAGE

Journal of Occupational and Organizational Psychology

Put me in coach: A daily examination of automated coaching on need for self-knowledge and learning goal orientation through metacognitive activities · 2025 · Learn More →

MSN

Employees are turning to AI for career advice. Cloverleaf’s research says that advice is: quit
· 2026 · Learn More →

News18

AI gives worse workplace advice to employees with less power: Report · 2026 · Learn More →