Five AI Agents, One Hidden Source: The Independence Illusion in Multi-Agent Decisions

Authors

  • Kwan Hong TAN Singapore University of Social Sciences, Singapore Author

Keywords:

multi-agent AI, correlated errors, consensus reliability, model diversity, AI assurance, human oversight

Abstract

Multi-agent artificial intelligence systems often convert agreement into confidence. That conversion is valid only when the agents contribute sufficiently independent evidence. This paper develops the independence illusion, defined as the overstatement of decision reliability that occurs when correlated AI outputs are counted as independent confirmations. A common-shock model derives the probability and reliability of unanimous agreement among n agents with marginal accuracy p and pairwise correctness correlation ρ. The model yields a counterintuitive result: positive correlation makes unanimity more frequent while making it less diagnostic. For five agents that are each 75 percent accurate, an independence-based calculation assigns 99.59 percent reliability to unanimity. At a correlation of 0.60, the correct reliability under the common-shock model is only 78.37 percent, even though unanimity occurs on 69.53 percent of tasks. A fully specified Monte Carlo design with 1,000,000 synthetic binary decisions then compares one-lineage, three-lineage, and five-lineage architectures. Five-vote majority accuracy rises from 80.8 percent in the one-lineage architecture to 84.8 percent with three lineages and 88.9 percent with five. Requiring an independent 90 percent accurate challenger to agree with a unanimous one-lineage panel raises accepted-case accuracy to 97.0 percent, but reduces automatic coverage to 50.5 percent. A beta-binomial robustness model preserves the central result. The paper proposes provenance-adjusted effective evidence, vote-pattern calibration, disagreement preservation, and independent challenge gates. The findings do not show that multi-agent systems are generally unreliable. They show that agent count is not evidence count, and that model lineage must enter assurance claims whenever consensus is used to automate consequential decisions.

0 0

References

1. Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). Artificial Intelligence Risk Management Framework Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology. [https://doi.org/10.6028/NIST.AI.600-1](https://doi.org/10.6028/NIST.AI.600-1)

2. Berg, S. (1993). Condorcet's jury theorem, dependency among jurors. *Social Choice and Welfare, 10*, 87–95. [https://doi.org/10.1007/BF00187435](https://doi.org/10.1007/BF00187435)

3. Boland, P. J. (1989). Majority systems and the Condorcet jury theorem. *Journal of the Royal Statistical Society Series D: The Statistician, 38*(3), 181–189. [https://doi.org/10.2307/2348873](https://doi.org/10.2307/2348873)

4. Breiman, L. (2001). Random forests. *Machine Learning, 45*, 5–32. [https://doi.org/10.1023/A:1010933404324](https://doi.org/10.1023/A:1010933404324)

5. Chen, Y., Niu, G., Cheng, J., Han, B., & Sugiyama, M. (2026). When and why does multi-agent debate fail and does it really underperform? *arXiv*. [https://doi.org/10.48550/arXiv.2510.20963](https://doi.org/10.48550/arXiv.2510.20963)

6. Clemen, R. T., & Winkler, R. L. (1999). Combining probability distributions from experts in risk analysis. *Risk Analysis, 19*(2), 187–203. [https://doi.org/10.1023/A:1006917509560](https://doi.org/10.1023/A:1006917509560)

7. DeGroot, M. H. (1974). Reaching a consensus. *Journal of the American Statistical Association, 69*(345), 118–121. [https://doi.org/10.2307/2285509](https://doi.org/10.2307/2285509)

8. Dietterich, T. G. (2000). Ensemble methods in machine learning. In J. Kittler & F. Roli (Eds.), *Multiple Classifier Systems* (Lecture Notes in Computer Science, Vol. 1857, pp. 1–15). Springer. [https://doi.org/10.1007/3-540-45014-9_1](https://doi.org/10.1007/3-540-45014-9_1)

9. Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. *Journal of Experimental Psychology: General, 144*(1), 114–126. [https://doi.org/10.1037/xge0000033](https://doi.org/10.1037/xge0000033)

10. Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., & Mordatch, I. (2023). Improving factuality and reasoning in language models through multiagent debate. *arXiv*. [https://doi.org/10.48550/arXiv.2305.14325](https://doi.org/10.48550/arXiv.2305.14325)

11. Geirhos, R., Jacobsen, J. H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., & Wichmann, F. A. (2020). Shortcut learning in deep neural networks. *Nature Machine Intelligence, 2*, 665–673. [https://doi.org/10.1038/s42256-020-00257-z](https://doi.org/10.1038/s42256-020-00257-z)

12. Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated errors in large language models. *arXiv preprint arXiv:2506.07962*. [https://doi.org/10.48550/arXiv.2506.07962](https://doi.org/10.48550/arXiv.2506.07962)

13. Kish, L. (1965). *Survey sampling.* John Wiley and Sons.

14. Kleinberg, J., & Raghavan, M. (2021). Algorithmic monoculture and social welfare. *Proceedings of the National Academy of Sciences, 118*(22), e2018340118. [https://doi.org/10.1073/pnas.2018340118](https://doi.org/10.1073/pnas.2018340118)

15. Kuncheva, L. I., & Whitaker, C. J. (2003). Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy. *Machine Learning, 51*, 181–207. [https://doi.org/10.1023/A:1022859003006](https://doi.org/10.1023/A:1022859003006)

16. Ladha, K. K. (1992). The Condorcet jury theorem, free speech, and correlated votes. *American Journal of Political Science, 36*(3), 617–634. [https://doi.org/10.2307/2111584](https://doi.org/10.2307/2111584)

17. Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. *Organizational Behavior and Human Decision Processes, 151*, 90–103. [https://doi.org/10.1016/j.obhdp.2018.12.005](https://doi.org/10.1016/j.obhdp.2018.12.005)

18. Lorenz, J., Rauhut, H., Schweitzer, F., & Helbing, D. (2011). How social influence can undermine the wisdom of crowd effect. *Proceedings of the National Academy of Sciences, 108*(22), 9020–9025. [https://doi.org/10.1073/pnas.1008636108](https://doi.org/10.1073/pnas.1008636108)

19. Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. *Nature, 627*, 49–58. [https://doi.org/10.1038/s41586-024-07146-0](https://doi.org/10.1038/s41586-024-07146-0)

20. Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. *Human Factors, 52*(3), 381–410. [https://doi.org/10.1177/0018720810376055](https://doi.org/10.1177/0018720810376055)

21. Smit, A. P., Grinsztajn, N., Duckworth, P., Barrett, T. D., & Pretorius, A. (2024). Should we be going MAD? A look at multi-agent debate strategies for LLMs. *Proceedings of the 41st International Conference on Machine Learning, 235*, 45883–45905. [https://proceedings.mlr.press/v235/smit24a.html](https://proceedings.mlr.press/v235/smit24a.html)

22. Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W., & Smyth, P. (2025). What large language models know and what people think they know. *Nature Machine Intelligence, 7*, 221–231. [https://doi.org/10.1038/s42256-024-00976-7](https://doi.org/10.1038/s42256-024-00976-7)

23. Tabassi, E. (2023). *Artificial Intelligence Risk Management Framework (AI RMF 1.0)* (NIST AI 100-1). National Institute of Standards and Technology. [https://doi.org/10.6028/NIST.AI.100-1](https://doi.org/10.6028/NIST.AI.100-1)

24. Tan, K. H. (2026a). The AI assurance cost paradox: Fixed governance costs, heterogeneous-firm adoption, and market concentration. *Journal of Advanced Multidisciplinary Studies, 1*(1), 310–329. [https://doi.org/10.68050/JAMS103](https://doi.org/10.68050/JAMS103)

25. Tan, K. H. (2026b). The coordination compression trap: A dynamic model of AI-enabled productivity, verification debt, and organisational resilience in global knowledge firms. *Journal of Global Economics Management and Business Research, 18*(3), 256–278. [https://doi.org/10.56557/jgembr/2026/v18i310956](https://doi.org/10.56557/jgembr/2026/v18i310956)

26. Tan, K. H. (2026c). Runtime assurance for enterprise agentic AI systems: A policy-gated control model with quantitative autonomy-risk scoring. *World Journal of Advanced Research and Reviews, 31*(1), 512–522. [https://doi.org/10.30574/wjarr.2026.31.1.1872](https://doi.org/10.30574/wjarr.2026.31.1.1872)

27. Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-consistency improves chain of thought reasoning in language models. *International Conference on Learning Representations*. [https://openreview.net/forum?id=1PL1NIMMrw](https://openreview.net/forum?id=1PL1NIMMrw)

Downloads

Published

2026-09-22

How to Cite

Five AI Agents, One Hidden Source: The Independence Illusion in Multi-Agent Decisions. (2026). Journal of Advanced Multidisciplinary Studies (JAMS), Page 164-178. https://jamsjournal.org/JAMS/article/view/489

Similar Articles

91-100 of 156

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)