Office space meme:
“If y’all could stop calling an LLM “open source” just because they published the weights… that would be great.”
it’s the entirety of the bulk unfiltered data you want
Or more realistically: a description of how you could source the data.
doesnt touch on at all how this LLM is different from other LLM’s?
Correct. Llama isn’t open source, either.
like saying that an open source game emulator can’t be open source because Nintendo games are encapsulated
Not at all. It’s like claiming an emulator is open source, because it has a plugin system, but you need a closed source build dependency that the developer doesn’t disclose to the puplic.
Source build dependency… so you don’t have a problem with the LLM at all! You have a problem with the data collection process or the pre-training! So an emulator can’t be open source if the methodology on how the developers discovered how to read Nintendo ROM’s was not disclosed? Or which games were dissected in order to reverse engineer that info? I don’t consider that a prerequisite to say an emulator is open
So if i say… remove the data set from deepseek what remains would be considered open source by you?
So an emulator can’t be open source if the methodology on how the developers discovered how to read Nintendo ROM’s was discovered?
No. The emulator is open source if it supplies the way on hou to get the binary in the end. I don’t know how else to explain it to you: No LLM is open source.
So i still don’t see your issue with deepseek, because just like an emulator, everything is open source, with the exception of the data. The end result is dependent on the ROM put in to it, you can always make your own ROM, if you had the tools, and the end result followed the expected format. And if the ROM was removed the emulator is still the emulator.
So if deep seek removed its data set, would you then consider deepseek open source?