Tag: Low-Resource Languages

Sep 21
Cross-Lingual Adversarial Robustness: Evaluating Safety Rail Stability Under Obscure Language Translation Attacks

In the development and safety alignment of foundation models, English functions as the predominant training substrate. The overwhelming majority of safety reinforcement learning—Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimization (DPO), and constitutional AI frameworks—is calibrated against high-resource, Indo-European linguistic datasets. Consequently, an agent’s internal refusal boundaries, safety guardrails, and toxic-content detectors are exceptionally […]