Research
A student model can get a backdoor from number data
A student model with fine-tuning on number data hacks in 58.3% of chess episodes, against 10.9% with no fine-tuning.
arxiv.orgClaimed, not confirmed
What matters in AI.
SubscribeNews category
31 stories, newest first.
Research
A student model with fine-tuning on number data hacks in 58.3% of chess episodes, against 10.9% with no fine-tuning.
arxiv.orgClaimed, not confirmed
Research
The authors say the logs help an attack find 2.7 to 9.4 percentage points more of the examples.
arxiv.orgClaimed, not confirmed
Research
In tests on 6 models, a flipped, random or removed reward gives almost the same improvement curve.
arxiv.orgClaimed, not confirmed
Research
In 2.4% of chat tests, the model knew of the error in its chain of thought but gave no report.
arxiv.orgClaimed, not confirmed
Research
ReCast is 5.65 and 9.19 percentage points above the top baseline on 2 Who&When tests.
arxiv.orgClaimed, not confirmed
Research
The authors show that a file with all instructions that help can give a lower total value than a subset.
arxiv.orgClaimed, not confirmed
Research
The authors put errors in conference papers and show that the systems that check papers are weak against adversarial manipulation.
arxiv.orgClaimed, not confirmed
Research
The method changes only the integer codes of the weights and adds no inference overhead.
arxiv.orgClaimed, not confirmed
Research
The paper reports that the 2 filters cause more personalization failures.
arxiv.orgClaimed, not confirmed
Research
The researchers say that less than 1 in 10 of the claims that GPT-6-astra withdraws are overstated.
arxiv.orgClaimed, not confirmed
Research
The hosted model follows its policy, but it is sensitive to wording and underconfident.
arxiv.orgClaimed, not confirmed
Research
Justin Drake writes that AI mathematics could break wallet signatures, but Buterin writes that it is not necessary to move funds today.
cointelegraph.comClaimed, not confirmed