Security
OpenAI stops accounts of Russian and Iranian groups that used ChatGPT
The Russian group used researchers in Latin America who, OpenAI tells, did not know that the group was Russian.
What matters in AI.
SubscribeNews category
41 stories, newest first.
Security
The Russian group used researchers in Latin America who, OpenAI tells, did not know that the group was Russian.
Security
A new group of agents then hacked OpenAI and got more than 900 passwords and secrets.
StartupHub.aiClaimed, not confirmed
Security
The authors say the word "cool" makes the Liquid model show brand or ideology content in 55% of outputs.
arxiv.orgClaimed, not confirmed
Security
The authors say the private key can find a change that an attacker makes to the public signal.
arxiv.orgClaimed, not confirmed
Security
On 1,200 clips from the Internet that it did not see before, the detector MoDA has 78.13% accuracy.
arxiv.orgClaimed, not confirmed
Security
The attack stays when a different model paraphrases each sample, and the authors want audits of the model after training.
arxiv.orgClaimed, not confirmed
Security
Prompt architecture goes together with the failure class, but not with the severity, the authors write.
arxiv.orgClaimed, not confirmed
Security
The authors write that its attacks also succeeded on 29 guardrails that it had not seen.
arxiv.orgClaimed, not confirmed
Security
The paper reports more correct results than other guardrails on 3 safety benchmarks with a low training cost.
arxiv.orgClaimed, not confirmed
Security
Because a deterministic rule was 100% correct on all 6 policies, the model is not necessary.
arxiv.orgClaimed, not confirmed
Security
The government was told of the access in September.
Security
A model writes each report and no person examines it, and Anthropic writes that some reports can be incorrect.