Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams

About

Team science holds that leadership is contingent: it helps only under specific conditions, and capable, autonomous teams may need none at all. We ask the analogous question for multi-agent LLM teams: under what measurable conditions does process-level coordination control add value, and do those conditions match what team science predicts? We use behavioral signatures (majority lock-in, exploration, recovery from an incorrect round-0 consensus) and per-action ablations, clean because each controller is an explicit action set, not a monolithic prompt. We operationalize three classical leadership styles (transactional, transformational, situational) as controllers over a shared action vocabulary (explore, revise, accept, synthesize). A matched controller with the same actions but an arbitrary rule recovers no better than majority voting, so the theory-derived rule, not the vocabulary, does the work. Across four task regimes and three open-weight model families, no controller dominates by accuracy, as the contingency view predicts: transactional control matches a shared round-0 vote on all 12 (model, regime) combinations to within 1.3pp, and gains appear only on the one combination where the round-0 majority is unreliable (llama-4-scout social; situational +8pp over flat). A recovery-advantage account, tested with four boundary probes, says a controller beats plain interaction only where the round-0 majority is unreliable, the task is recoverable, and undirected interaction does not already repair it. These regions map onto contingency theory (leadership substitutes, path-goal redundancy, the situational readiness gap), so a largely null accuracy result is what the theory predicts, not a failure of the controllers. We read process-level coordination control as a contingency to be measured and theory-mapped, not a leaderboard to be topped.

Haewoon Kwak• 2026

Related benchmarks

TaskDatasetResultRank
Closed-ended reasoningClosed regime n = 300
Exact Match Accuracy87.9
21
Social ReasoningSocial regime (n = 300)
Exact Match Accuracy51.3
21
Abductive Natural Language InferenceAlphaNLI n=300 per condition (test)
Accuracy93
21
Mixed reasoning tasksMixed workload n = 450
Exact-match Accuracy73.6
21
Multi-agent DebateSocial regime
Accuracy51.3
15
Multi-agent DebateMixed regime
Accuracy71.8
15
Mathematical Problem SolvingMATH-500 simple-numeric-answer (Level 5)
Accuracy93.6
15
Multi-agent DebateClosed regime
Accuracy85.7
15
Multi-agent DebateAlphaNLI
Accuracy92.7
15
Showing 9 of 9 rows

Other info

Follow for update