Phronesis Activation-Steering Failure-Mode Labels (FM-X)
收藏资源简介:
A corpus of ~2,966 language-model generations from activation-steering experiments, each read in full and labeled by an LLM judge (Anthropic Claude, Opus-family) under a frozen, human-authored protocol with author review — no regex or automatic scoring was used for any verdict. Each generation carries a verdict plus, on failure, a tagged failure mode from a 13+ category taxonomy (FM-1..FM-13). It documents where automatic/regex scorers systematically mislabel steered LLM output — both false positives (crediting degenerate or confabulated text) and false negatives (missing genuine behaviour in non-standard prose). Three CSV batteries span the qwen2.5 / qwen3 / deepseek-r1-distill / phi-4 / llama-3.1 families. Useful as a study set for LLM-as-judge robustness and a catalogue of steered-LLM failure modes. Note for LLM-as-judge research: the labels are themselves LLM-judge outputs — evaluating a judge (especially a Claude-family one) against this dataset partly measures consistency with Claude under this protocol, not independent human ground truth. AI is not a listed author; any errors are the author's. See README.md for the full taxonomy, schema, and labelling protocol.



