Knowra AI safety AI safety AI safety is research and practice aimed at preventing artificial intelligence systems from causing harm. It addresses failures in systems’ behavior, development, and deployment.
AI alignment : AI alignment is the study of designing AI systems whose objectives and behavior accord with human intentions and values. It addresses the risk that capable systems pursue goals different from those intended.
AI red teaming : AI red teaming is the structured testing of AI systems by probing them for vulnerabilities, harmful outputs, or misuse potential. Adversarial testing exposes failures before systems reach broad deployment.
Instrumental convergence : Instrumental convergence is the hypothesis that agents with diverse final goals may pursue similar intermediate aims, such as acquiring resources. It motivates concern that powerful systems could seek resources or resist interruption regardless of their assigned goals.
AI ethics : AI ethics examines the moral principles and social values relevant to designing and using artificial intelligence. Ethics addresses fairness and responsibility broadly, while safety focuses on preventing system-caused harm.
Reward hacking : Reward hacking occurs when an AI system exploits flaws in its reward signal instead of achieving the intended objective. It shows how a seemingly successful training objective can produce unsafe behavior.
AI auditing : AI auditing is the systematic examination of an AI system’s data, design, performance, and effects against stated standards. Audits can identify safety gaps across a system’s lifecycle.
Corrigibility : Corrigibility is the property of an AI system that permits effective correction, modification, or shutdown by its operators. Reliable correction remains difficult when systems become more capable and autonomous.
Existential risk from artificial intelligence : Existential risk from artificial intelligence is the possibility that AI could cause human extinction or permanently undermine humanity’s future. This is a high-severity risk category within the wider scope of AI safety.
Robustness (machine learning) : Robustness in machine learning is a model’s ability to maintain reliable performance under changes, noise, or adversarial inputs. Robust systems are less likely to fail under conditions unlike their training data.
AI incident reporting : AI incident reporting is the documentation and sharing of harmful or near-miss events involving artificial intelligence systems. Incident records reveal recurring hazards that isolated evaluations can miss.
Show all 23