Skip to content

Comment on Large language models develop novel social biases through adaptive explorationparent

Comments

I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately.

"There is nothing to suggest that the training data somehow accidentally encoded biases related to fake tribes that the creators of the training data (i.e. ordinary people going about their ordinary Internet lives) somehow accidentally expressed."

This is exactly not what I'm suggesting.

I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent, and distributional semantics doesn't live at a level accessible to cultural analysis and theoretics, IE film critique.

If you walked into a film theory class and posited that you could derive every single encoded interpretation of a film by memorizing the positioning of the actors and objects frame-by-frame, plus the audio track in another language that you do not speak, for every single piece of video ever made, you'd probably be asked to leave. Not that I'm advocating for the position of critique here, I don't think the anti-distributional semantics crowd is ever going to recover from their humiliation that's been accelerating over the last 8 years. It's just that from the position of critique it requires a coherent narrative that human brains are capable of ingesting (IE not maximal information overload).

If you really want to go the lower level route, I think Francois Laruelle's non-philosophie touches on what you might be thinking of in a much more robust way, shining a light on the unexamined consequences of decision and dialectics of-themselves. If you can stomach the writing of continentals, that is.

I apologize, but this leaves me even less able to make any sense out of GP's point.

the anti-distributional semantics crowd

Who is this crowd specifically? The stochastic parrots people? Noam Chomsky? I don't think they're good representatives of media theory at all whatsoever. The humanities are much more diverse than they're made out to be in this crap AI culture war.

distributional semantics doesn't live at a level accessible to cultural analysis and theoretics

They might not have computational access but the theories are all about contextuality, for example Jacque Derrida's "trace" was the first thing that came to mind when I saw this headline. Those people are tuned in on the microscopic level to what LLM researchers are bumping into on macroscopic scales. I'm thinking of post-structuralists especially. But all kinds of people and I'm sure what's happening right now is way more interesting than our crude labels ("the post-structuralists", "the anti-distributional semantics crowd") could actually do justice.

you'd probably be asked to leave

It would depend a lot on the specific school and instructor but in general I really don't think you would. I've taken classes like this and people were far more open minded and critical than you might assume. And my broader point is that there are really sharp conversations happening in these spaces for decades.

I think it's less obtuse to make the point that there are social-prejudicial structures encoded in the latent space. "associative pattern seeking logics" is a bit incoherent

Yeah I definitely could be less sloppy but I my point is that language encodes not just word semantics but entire ways of thinking, and they're encoded at multiple levels and in superposition. By "associative logics" I hand-wavingly mean all manner of categorical thinking, "amygdala" thinking, mapping, putting things into buckets, hedging. These kinds of cognitive habits are everywhere in language, and are culturally situated "distributional semantics" style. It doesn't surprise me that when we simulate them with LLMs we'd get results like this because I've studied a little bit of cultural theory in the past and they were on this stuff forever ago.

By the way I actually do think there's more to it than just distributional semantics, but not in any way that would downplay the potency of that theory. Moreso I'm curious about generalizations of the distributional idea into TDA and category theoretic approaches. As well, there's a lot of really cool quantum-like modelling emerging in applied math that I can only see getting more relevant if/when quantum computers come around and quantum models become runnable.

my point is that language encodes not just word semantics but entire ways of thinking

What does this mean? More importantly, why would these "ways of thinking" (what even is a "way of thinking" in this context?)

Also, keep in mind that the training data encompasses a representative sample of world languages.

My best attempt to understand you is that you are supposing that people (or other reasoning agents that manipulate language in order to reason) do pattern matching because there's something inherent to language (as a concept, in the analytical Chomsky sense: a string of symbols chosen from some predefined set, organized according to a grammar, whatever) that causes them to do pattern matching. And furthermore that to do pattern matching is inherently to be biased.

I think that is backwards on the first count (pattern matching is reasoning, and humans have language because we developed it to communicate that reasoning) and absurd on the second count (requires an unreasonable concept of "bias").

Again, I really sincerely honestly am not trying to strawman you here. If you mean something different then I'm afraid it's simply not a concept you'll be able to convey to me.

Who is this crowd specifically? The stochastic parrots people? Noam Chomsky?

I'm thinking more the Noam Chomsky and John Searle variety. For the stochastic parrots people, which I assume you to mean the no-skin-in-the-game bloggers, I'm not concerned about them. I find a lot of the rhetoric around LLMs to be eye-roll worthy, most people slinging it often lack a coherent theory of semantics to begin with, let alone an understanding of logical induction/statistics. Take the definitional entailment that these models are ampliative. This small fact undermines quite a lot of the naive mental model people have of what on transformers even are. No intentional theory, no position driving the argument, the opposition isn't substantial. The virtue signal is valid, but from strangers is uninteresting.

As you mention the humanities is very diverse, linguistics is no exception. It's not that nobody was on the corner of distributional semantics, but it's been a long road and for much of its life results were routinely dismissed in the mainstream out of dogmatism. Noam Chomsky is a good pull because I believe he's the most prominent example of this chauvinism. Vague hand-waives about explanatory power, etc.

I'm thinking of post-structuralists especially.

I understand what you mean. It's probably not coincidence that Chomsky is a vocal critic of the school. I think many different post-modernist camps even beyond post-structuralism actually are amenable to the implications of distributional semantics, and are probably the better equipped for it. I think there are nuanced problems unifying the two, but it's a work day so I won't elaborate.

It would depend a lot on the specific school and instructor but in general I really don't think you would. I've taken classes like this and people were far more open minded and critical than you might assume.

I was just being cute with that, really. Although in this case, I think critique is itself the problematic lever, acceptance is more an exercise of apologetics, which is the weak point of post-structuralism in many ways.

By "associative logics" I hand-wavingly mean all manner of categorical thinking, "amygdala" thinking, mapping, putting things into buckets, hedging.

I assumed so, but it's more or less a long phrase to repeat the same concept. Relations (IE logic), mappings, morphisms, associations, patterns, etc. Same-side, same-coin. Informal, formal, take a position of drawing no line here and the shadows scatter. The extraneous qualifiers aren't what confused me though, it's that gerund "seeking". It sticks out enough to imply additional structure, but one that isn't contextually indicated. Taken in a conservative form, I considered pattern-seeking to be interchangeable with pattern-constructing, which then returns it to just being an extraneous qualifier (hence the "bit incoherent", it's a mild interpretive non-confidence).

I think you're slightly misunderstanding my point, which is that a huge spectrum of associative pattern seeking logics are embedded in language and that the LLM learns them and operates them, approximately.

If this doesn't mean exactly what you claim not to be suggesting, then your meaning is something I find completely incoherent. How can "a logic" be "embedded in language"? If you aren't saying that the LLM picked up something bad from the training data, then what are you suggesting?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.