Algorithmic Cooperation: A Comparison with Human Play in the Infinitely Repeated Prisoner's Dilemma

Reinforcement learning algorithms play an increasingly important role in economic situations. These situations are often strategic, and the artificial intelligence may or may not be cooperative. We compare human and algorithmic cooperation rates in the infinitely repeated two-player prisoner's dilemma and study which strategies they choose to cooperate and punish deviations. Through a sequence of computational Q-learning and human-player experiments, we find that our Q-learning algorithms tend to cooperate less than humans, particularly when cooperation is risky or not incentive-compatible. Algorithms often use different strategies than humans, leading to distinct on- and off-path behavior.