Prompt clarity and quality
Assess how clearly the engineer states the goal and supplies useful context and constraints.
Agent Effectiveness analyzes coding-agent sessions to help teams understand which practices produce useful, verifiable work. Compare session patterns, review coaching suggestions, and share practices across the team.
Claude CodeCodexCursorOpenCodeGitHub Copilot
Spotlights
Example sessions · Last 30 days
74 / 100
team effectiveness, example top quartile 81
12 of 64
engineers already scoring 80 or above
Plan before you build
19 engineers open feature work with a plan request. 34% fewer iterations, 6 example prompts.
+4 Feature development$1,900 / mo
Adoption 19 of 64 · discuss as a team practice
Name the files before the agent edits
Diffs 2.2× smaller when the prompt names the edit surface. 23% of feature sessions do.
+3 Feature development
Adoption 14 of 64
Run the suite before reporting done
56 of 97 bug-fixing sessions already do. The other 41 needed 2.9× more follow-up fixes.
Suggested CLAUDE.md practice+5 Bug fixing
Adoption 56 of 97 sessions · review the guide change
Illustrative practice suggestions. Confirm access and attribution settings with your organization's administrator.
Trusted by AI-forward enterprises
Agent mix
Larridin brings coding-agent sessions, work patterns, token usage, and cost into one view. Session cost estimates use model pricing and usage telemetry; provider billing views show billed spend where a billing connection is available.
Agent mix
$21,480
estimated cost of analyzed sessions
612
analyzed sessions of 1,529 captured
$35
average per session, median $4
28%
of estimated cost in six sessions
Daily estimated cost by type of work
weekends lighter
Prompt cache
input tokens, all agents
91%
of input tokens read from cache in this example. Cache share varies across sessions.
$9,240
estimated difference from uncached input pricing in this example.
$412
estimated cache-creation cost across 62 example sessions.
Illustrative cost estimates use model list rates, including separate cache-creation and cache-read prices. Estimates can differ from invoices. Scores summarize the analyzed sessions in each group.
Sessions
Review available session traces, agent turns, token usage, estimated cost, and linked PRs. The score combines six dimensions: prompt clarity, session steering, sophistication, prompt quality, verification discipline, and task outcomes. Available dimensions carry equal weight.
Agent Traces
App: AllPrompts
32
30h 23m open
Engaged
2h 40m
4 re-warms
Cost estimate
$37.54
Example model mix
Tokens
76.5M
99.5% cached
Overall score
A strong session. The engineer opened with an explicit goal and asked for a plan, held scope with five targeted corrections, and the agent verified its work before reporting completion.
Show full reasoning →
16 citations · 9 prompts, 7 agent turns
Selected dimensions · 3 of 6
team average
Asks are scoped with an explicit session goal, constraints and an upfront plan request.
Each prompt combines the goal, a performance target and constraints for the CI speed-ups.
Catches drift quickly with targeted corrections and clear scope boundaries.
Scoring
Session scores bring six equally weighted dimensions into a team view of agent use. Compare practices, inspect the sessions behind a pattern, and turn useful findings into guidance your team can review and adopt.
Effectiveness
74 / 100
▲ 2 vs August average of 72
12 of 64
engineers scoring 80 or above
6
dimensions in the effectiveness score
Working well
Plan before you build
19 engineers · feature sessions finish in 34% fewer iterations
Name the files before the agent edits
23% of feature sessions · diffs 2.2× smaller
Targeted corrections instead of restarts
41 engineers · steering at top-quartile level
Tests run before reporting done
56 of 97 bug-fixing sessions · the fewest follow-up fixes in the org
Needs attention
Closing without a test run
41 of 97 bug-fixing sessions · 2.9× more follow-up fixes
Opening with "fix the bug"
33 sessions · 12 turns before a plan appears
Restarting instead of steering
27 sessions · a fresh session after each miss, context rebuilt every time
Idle gaps that rebuild the cache
62 sessions · $412 in cache rebuilds, no score effect
The score averages available ratings for prompt clarity, session steering, sophistication, prompt quality, verification discipline, and task outcomes, scaled to 100. Comparisons here are illustrative.
Example practice guidance
review with your team
Run the test suite before reporting done
41 sessions closed without one, then needed 2.9× more follow-up fixes.
CLAUDE.md · apps/api+5 Bug fixing$640 / mo
Ask for a plan before feature work starts
19 of 64 engineers already do. Share their pattern with the other 45.
Spotlight+4 Feature development$1,900 / mo
Compare models for codebase research
96 research sessions. A lower-cost model scored 78 against 79 at 42% less in this example.
Model comparison$900 / mo
Example guide change · apps/api/CLAUDE.md
For you
example coachingSix of your nine bug-fixing sessions this month closed without a test run. Teammates who run the suite first finish with 2.9× fewer follow-ups. Try /verify on the next one.
Example coaching based on session patterns. Access depends on your organization's roles and configuration.
Privacy
Agent sessions can contain sensitive context. Larridin scopes analytics by organization and role. Work with your administrator to confirm which session details are collected, who can access them, and the retention settings for your deployment.
Access planning example
This table illustrates an access policy to review with your administrator. It does not describe default permissions; available controls depend on your deployment.
A closer look
Each scored session is evaluated across six skills. The composite is the mean of the available 1–5 dimension scores, converted to a 0–100 scale. Unobserved dimensions are excluded rather than treated as zero.
Assess how clearly the engineer states the goal and supplies useful context and constraints.
Assess how the engineer guides the session and uses the agent’s capabilities to work through the task.
Assess whether the work is checked and what the session accomplished. Follow the session evidence behind the score.
| Coding agents | What to check before comparing |
|---|---|
| Claude Code, Codex, Cursor, OpenCode | Confirm session capture and attribution for the versions your organization uses. |
| GitHub Copilot, Devin, Gemini CLI, Antigravity | A tool appearing in the catalog does not imply identical session coverage. Review captured and unscored sessions. |
It measures observed skills in a coding-agent session: prompt clarity, session steering, sophistication, prompt quality, verification discipline, and task outcomes. It is not an overall rating of an engineer’s productivity.
Compare sessions with similar work types and coverage using the same rubric. Read the evidence and number of scored sessions alongside any average. Differences can reflect task difficulty, capture coverage, or work mix.
Session data is organization-scoped and access depends on role, enabled tools, and sharing configuration. Review access controls with your administrator before capturing sensitive sessions. Model-provider retention and Larridin session storage are separate controls.
The product views on this page use illustrative data to explain the metrics and workflows. For metric definitions, sample calculations, and assumptions, see our measurement methodology.
Connect your coding agents and review the sessions available to your team. Try it on your own repos, or get pricing for your org.