Final Unicourse'tan Çalış, Yüksek Notu Garantile!
Vizesine Unicourse'tan Çalış, Yüksek Notu Garantile!
Someone should make a community to freely distribute examples of data poisoning people can randomly put in their social media posts/images to sabotage AI
-
This post did not contain any content.
-
This post did not contain any content.
Na, for it to be effective it needs to be wide spread, but if its wide spread then it can be filtered out of the training material.
-
Na, for it to be effective it needs to be wide spread, but if its wide spread then it can be filtered out of the training material.
I've read in papers that you can poison datasets with a very small percentage of the data, if done cleverly. I can fish up the source if you want (but it might take me some time).
edit: here it is.
We conduct the largest pretraining poisoning experiments to date, pretraining models from 600M to 13B parameters on chinchilla-optimal datasets (6B to 260B tokens). We find that 250 poisoned documents similarly compromise models across all model and dataset sizes (...)
Emphasis mine. All it takes is 250 poisoned documents.
-
Na, for it to be effective it needs to be wide spread, but if its wide spread then it can be filtered out of the training material.
There's a new technique that uses the AIs "thinking" tags to get it to do things that are otherwise banned by policy.
I'll have to find the article again. But due to the way LLMs work, they can't defend against this sort of attack.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Kayıt Ol Giriş