Small Foundation Models of Human Cognition and Behaviour - CloudYume
Small Foundation Models of Human Cognition and Behaviour
· /
사실 흐름
arXiv:2608.05224v1 Announce Type: new Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process
从更广阔的视角看,这类研究正推动人工智能与神经科学的深度融合。有学者在《人类神经科学前沿》期刊撰文指出,传统训练数据依赖人类认知的行为输出,而神经成像数据则能打开一扇窗,让我们窥见产生这些行为的潜在认知过程,从而可能提升模型的鲁棒性和通用性[来源:Frontiers in Human Neuroscience]。将行为数据与脑数据结合,优先用于基础模型训练中传统数据局限性最明显的高价值步骤,是一条充满前景的路径。这或许能帮助我们构建不仅模仿人类说什么、写什么,更能模拟人类如何思考、如何决策的模型。当然,这些模型目前仍是用于科研的计算工具,任何将其应用于临床或心理评估的尝试,都需经过专业人员的严格检验。
arXiv:2608.05224v1 Announce Type: new Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.