New book publication!

The Book cover for Artificial Knowledge of Language: A Linguist's Perspective on Its Nature, Origins and Use

Almost three years ago I wrote a post responding to some early LLM hype, and somewhat unexpectedly, that post served as the kernel for a chapter in a book that has just been published! The book, Artificial Knowledge of Language: A Linguist’s Perspective on Its Nature, Origins and Use (Vernon Press) includes 8 essays from scholars offering a critical view of the relation between LLMs and the human Language Faculty.

On a personal note, the process of converting my post to a published chapter required me to more fully engage with the scholarly community around LLMs both in the published (usually not peer-reviewed) literature and undergoing anonymous peer review, and this engagement left me somewhat surprised at the intellectual bankruptcy in that community, and frightened that they are taken seriously by society at large.

Artificial Knowledge of Language: A Linguist’s Perspective on Its Nature, Origins and Use is edited by José-Luis Mendívil-Giró and published by Vernon Press

New (draft) paper: “FormSet and Parallel Derivation: a synthesis”

I have posted a draft of a new paper on LingBuzz. It’s a continuation of my 2022 Biolinguistics article, though I frame it as a critical analysis of certain recent developments in linguistic theory. The abstract is as follows:

This paper analyzes the FormSet operation as it has been defined and used in the literature specifically in the domain of conjunction structures. It finds that, though its use as a structure building operation must either contradict its definition and justification as a general-purpose preparatory operation or lead to unwanted empirical predictions. It goes on to show that, if we redefine MERGE as applying not to syntactic objects in a workspace but to n-ary sets of syntactic objects in a workspace, that we can retain the justification of FormSet without sacrificing empirical validity. Furthermore, it presents novel empirical arguments in favour of this redefined MERGE—that it correctly predicts the ambiguities raised by conjunction structure, that it correctly predicts a cross-linguistic gap in the plural anaphor system, and that it gives some insight into the system of number “features.”

As always I’m happy to hear any comments you have.

Chris Collins on “foundational work” in syntactic theory

Over on his blog, Chris Collins has a new post on the difference between what he calls the “Standard Paradigm for Syntactic Research” and the “Foundational Paradigm for Syntactic Research.” In it, he identifies a sociological phenomenon within generative syntax—there is a general apathy towards “foundational” work in the field—and suggests four explanations for it. I mostly agree with what he writes, but I wanted to highlight and deepen one of his explanations and add another to the mix.

Collins’ first explanation is as follows (emphasis mine):

[F]oundational work takes a lot of time and effort, and the payoff is uncertain. The question is what counts as progress. You may think for years and years about the ‘copy versus repetition’ distinction or about the definition of c-command or about the nature of workspaces without resolving the issues, even though in the end you have a much deeper understanding. Does that count as progress? Can you write it up and publish it? Does it count as currency in the academic monetary system?

I would posit that almost any “non-foundational” work can yield some sort of result, be that a new generalization, new data, or maybe a new synthesis of data, whereas “foundational” work does sometimes end in blind alleys with little to show for it. While this difference is important when measured in papers, posters, and talks, its also important when measured in money.

In most of the global north, public funding for the science, social science, the humanities, and the arts is distributed by members of the particular fields, ostensibly on an apolitical basis. Yet for almost as long as modern states have been funding science, social science, the humanities, and the arts there has been a faction within those states that have categorically opposed it. And one of the favoured rhetorical tactics of that faction is to trot out obscure studies that seem frivolous out of context and ask disingenuously “You think this is a good thing to fund??”

It is difficult to believe that anyone on a grant committee is not aware of this possibility, and has in the back of their mind “how would I explain giving this grant to a ‘foundational’ project that may yeild nothing?”

The explanation I’d add to Collins’ is a psychological one. What he calls “foundational” work in syntactic theory is theoretical work while “non-foundational” work in syntactic theory is actually empirical/analytical/experimental work, not theoretical work (cf Chametzky 1996). Most syntacticians doing pen-and-paper analysis, then, consider they’re work to be theoretical syntax—they see themselves as theoretical syntacticians. So when they encounter “foundational” work (i.e., theoretical work) it evokes some cognitive dissonance (“I know theoretical syntax, and this is supposed to be theoretical syntax, but I can’t follow it, or I don’t see the point.”) which often leads to hostile reactions. Indeed, suggesting to someone they’re not doing theoretical work is often taken as an insult due to the esteem given to “theoretical work.”

I’ve arrived at this explanation because, when I was in grad school I experienced the same cognitive dissonance form the other side. I was always interested in “foundational” syntactic theory, but found I couldn’t do or get excited about the the “theoretical” work that my syntax colleagues were doing. I simply don’t have the mind for the empirical details. The natural response to this was impostor’s syndrome, until I realized that what we’d been calling “theoretical syntax” was actually two very different research paradigms, and that the work that I couldn’t do, was a different type of work. This saved my self-esteem, but still left me with the understanding that there was a ceiling on my academic career, not because of my skills, but because of the state of the field.

New paper: On Modern Language Models, Impossible Languages, and Anti-science

I’ve just submitted a new paper to appear in a forthcoming volume edited by José-Luis Mendívil-Giró. The volume is on the implications of Large Language Models for linguistic theory and my paper is a refinement of what I’ve written here, here, and here. The abstract is:

While “modern language models,” which put into practice empiricist theories of language, are claimed to be refutations of rationalist theories of language, a close look at the claims made in their favour reveals otherwise. In this chapter, I critically review Piantadosi’s (2023) arguments for his claim that MLMs refute rationalist theories of language along with his replies to critiques of those arguments. I argue that, throughout his paper, Piantadosi misstates the claims, predictions, and arguments of contemporary rationalist theories of language, paying particular attention to the problem of impossible languages. I further argue that the goals of rationalist theories of language and MLMs are orthogonal to each other, with the former being a scientific inquiry aimed at explanation and understanding and the latter being an engineering project aimed at making tools, and that confusing the two, as Piantadosi does, can lead one to inadvertently take up an anti-scientific position.

If you’d like a preview, I’ve uploaded a draft on LingBuzz.

The DP Hypothesis—a case study of a sticky idea

Recently, in service of a course I’m teaching, I had a chance to revisit and fully engage with what might be the stickiest idea in generative syntax—The DP hypothesis. For those of you who aren’t linguists, the DP hypothesis, though highly technical, is fairly simple to get the gist of based on a couple of observations:

Observation 1: Words in sentences naturally cluster together into phrases like “the toys”, “to the store”, or “eat an apple.”

Observation 2: In every phrase, there is a single main word called the head of the phrase. So, for instance, the head of the phrase “eat an apple” is the verb “eat.”

These observations are formalized in syntactic theory, so that “eat an apple” is labeled a VP (Verb Phrase), while “to the store” is a PP (Preposition Phrase). Which leads us to the DP hypothesis: Phrases like “the toys,” “a red phone,” or “my dog” should be labelled as DPs (Determiner Phrases) because their heads are “the,” “a,” and “my,” which are called determiners in modern generative syntax.

This is fairly counterintuitive, to say the least. The intuitive hypothesis—the one that pretty much every linguist accepted until the 1980s—is that those phrases are NPs (Noun Phrases), but if we only accepted intuitive proposals, there’d be no science to speak of. Indeed, the all the good scientific theories start off counterintuitive and become intuitive only by force of argument. One of the joys of theory is experiencing that shift of mind-set—it can feel like magic when done right.

So it was quite unnerving when I started reading the actual arguments for the DP hypothesis, which I had, at one point, fully bought into, and and began to become less convinced by each one. It didn’t feel like magic, it felt like a con.

My source for this is a handbook chapter by Judy Bernstein that summarizes the basic argument for the DP Hypothesis—a twofold argument consisting of a Parallelism argument and purported direct evidence of the DP Hypothesis— as previously advanced sand developed by Szabolcsi, Abney, Longobardi, Kayne, Bernstein herself, and others.

The parallelism argument is based on another counterintuitive theory developed in in the mid-20th century which states that clauses, previously considered either headless or VPs, are actually headed by abstract (i.e., silent) words. That is, they are variously considered TPs (Tense Phrases), IP’s (Inflection Phrases), or CPs (Complementizer Phrases). The parallelism argument states that “if clauses are like that, then ‘noun phrases’ be like that too” and then finds data where “noun phrases” look like clauses in some way. This might seem reasonable on its face, but it’s a complete non sequitur. Maybe the structure of a “noun phrase” parallels that of a clause, but maybe it doesn’t. In fact, there’s probably good reason to think that the structure of “noun phrases” is the inverse of the structure of the clause—the clause “projects” from the verb, and verbs and nouns are complementary, so shouldn’t the noun have complementary properties to the verb?

Following through on parallelism, if extended VPs are actually CPs, then extended NPs are DPs. Once you have that hypothesis, you can start making “predictions” and checking if the data supports them. And of course there is data that becomes easy to explain once we have the DP Hypothesis. Again, this is good as far as it goes, but there’s a key word missing—”only.” We need data that only becomes easy to explain once we have the DP Hypothesis. And while I don’t have competing analyses for the data adduced for the DP Hypothesis at the ready—though Ben Bruening has one for at least one such phenomenon—I’m not really convinced that none exist.

And that’s the foundation of the DP Hypothesis, a weak argument resting on another weak argument. Yet, it’s a sticky one—I can count on one hand the contemporary generative syntacticians that have expressed skepticism about it. Why is it so sticky? My hypothesis is that it’s useful as a shibboleth and as a “project pump”.

Its usefulness as a shibboleth is fairly straightforward—there’s no quicker way to mark yourself as a generative syntactician than to put DPs in your tree diagrams. Even I find it jarring to see NPs in trees.

To see the utility of the DP Hypothesis as a “project pump”, one need only to look at the Cartography/Nanosyntax literature. Once you open up a space for invisible functional heads between N and D, you seem to find them everywhere. This, I think, is what Chomsky meant when he described the DP Hypothesis as “…very fruitful, leading to a lot of interesting
work” before saying “I’ve never really been convinced by it.” Who cares if it’s correct, it contains infinite dissertations!

Now maybe I’m being to hard on the DP and its fans. After all, as far as theoretical avenues go, the DP Hypothesis is something of a cul de sac, albeit a large one—the core theory doesn’t really care whether “the bee” is a DP or and NP, so what’s the harm? I could point out that by making such a feeble hypothesis our standard, we’ve opened ourselves to being dunked on my anti-generativists. Or I could bore you with such Romantic notions as “calling all things by their right names.” Instead, I’ll be practical and point out that, contrary to contemporary digital wisdom, the world is not infinite, and every bit of real estate given to the DP cul-de-sac in the form of journal articles, conference presentations, tenure-track hires, etc. is space that could be used otherwise. And, to torture the metaphor further, shouldn’t we try to use our real estate for work with a stronger foundation?

The “science” of modern “AI”

(or Piantadosi and MLMs again (II)—continuation of this post)

In my critique of Prof. Piantadosi’s manuscript “Modern language models refute Chomsky’s approach to language,” I point out that regardless of the respective empirical results of Generative Linguistics and MLMs, the latter does not supersede the former because the two have fundamentally different goals. Generative Linguistics aims to provide a rational explanation of a natural phenomenon while MLM are designed to simulate human language use. Piantadosi does not dispute this, but rather states that

… there is an interesting debate about the nature of science lurking here. The critics’ position seems to be that in order for something to be a scientific theory, it must be intuitively comprehensible to us. I disagree because there are many phenomena in nature which probably will never admit a
simple enough description for us to comprehend. We cannot just exclude these things from scientific inquiry.

p37 of v7 (emphasis in original)

Being one of the “critics” referred to here, I can grant the professor’s description of my position as basically accurate if a bit glib. But what is his position? He doesn’t say precisely, but we can make some inferences. In lieu of a clear statement of his position, for instance, Piantadosi follows the above quote with this:

There probably is no simple theory of a stock market (why IBM takes on a particular value) or dynamics in complex systems (why an O2 molecule hits a particular place on my eyeball). Certainly there are local, proximate causes (Tom Jones bid $142 for IBM; the O2 molecule was bumped by another), but when you start to trace these causes back into the complex system, you will quickly exceed our ability to understand the complex network of interactions.

p37 of v7

These are slightly bizarre comments, as we do have comprehensible (i.e., simple) theories of stock markets—the efficient markets hypothesis, for instance[1]This should not be taken as an endorsement of the efficient markets hypothesis—or any part of (neo)classical economics—as correct. A theory’s scientific-ness is no guarantee of its … Continue reading—and gases—the kinetic theory, for instance—which can give approximate predictions regarding real life events like the examples given. The professor’s view can be narrowed down slightly based onhis assertion that Rawski & Baumont (2023) “seem to misunderstand the linkage between experiment and theory” (p34 of v7)[2]This is a bold claim for Piantadosi to make given that he is a psychologist, while Lucie Baumont—the latter half of Rawski & Baumont—is an empirical astrophysicist. when they state that “Explanatory power, not predictive adequacy, forms the core of physics and ultimately all modern science.” (Rawski & Baumont 2023) It would seem clear, then, that, for Piantadosi at least, that a “theory” is scientific only insofar as it has predictive power.

This may seem like a reasonable characterization—despite myriad insinuations to the contrary, virtually no one believes that predictive power is unimportant—but as soon as one attempts to develop that characterization things get dicey. What, for instance, is the required level of accuracy and precision for science? And What sort of things should a true science be able to predict? To use one of Piantadosi’s examples, individual molecules are the primitives of the kinetic theory of gasses, and the theory makes precise predictions about the behaviour of a gas—i.e., gas molecules in aggregate—but it is highly doubtful that it would make predictions about the actual motion of a particular molecule in any situation. Surely, this would be too much to ask of any theory of physics, yet Piantadosi seems to believe it is within the realm of scientific inquiry.

There’s also a question of what it means to “predict” something. Piantadosi’s argument boils down to “MLMs are better than Chomsky’s approach theory, because they make more correct predictions,” yet nowhere does he explicitly say what those predictions are, nor does he document any tests of those predictions. Instead, we are treated to his prompts to a chatbot followed by the chatbot’s response. Perhaps these are the predictions. Perhaps they predict how a human would respond to such prompts. If so, then so much the worse for MLMs qua scientific theories because, even if MLMs were indistinguishable from humans, the odds of any two humans answering a single question the same way is vanishingly slim, and any way to determine a general similarity between utterances would almost certainly be either arbitrary or dependent on some theoretical framework. At best, MLMs simulate human language use, meaning they no more predict facts of language than a compass predicts facts of geometry.

Chomsky’s approach to theories of language, on the other hand makes clear predictions if one bothers to engage with it. The predictions are of the form “Given theoretical statement T, a competent speaker of language L will judge expression S as (un)acceptable in context C.” This is exactly the sort of prediction that one finds in other sciences—”if one performs precisely this action under precisely these conditions, one will observe precisely this reaction”—and the sort of prediction that is absent in Piantadosi’s paper.

Indeed these predictions seem to be absent in the entire contemporary “AI” discourse, and with good reason—”AI” is not a scientific enterprise. It’s an engineering project. A fact that is immediately obvious when one considers how it measures success—against a battery of predetermined arbitrary tests. MLM researchers, then, aren’t discovering truths, they’re building tools to spec, like good engineers.

This is not to cast aspersions on engineers, but it does raise a question—the core question: How exactly can an engineering project like MLMs refute a scientific theory like Generative Grammar?

Notes

Notes
↑1 This should not be taken as an endorsement of the efficient markets hypothesis—or any part of (neo)classical economics—as correct. A theory’s scientific-ness is no guarantee of its correctness.
↑2 This is a bold claim for Piantadosi to make given that he is a psychologist, while Lucie Baumont—the latter half of Rawski & Baumont—is an empirical astrophysicist.

Piantadosi and MLMs again (I)

Last spring, Steven Piantadosi, professor of psychology and neuroscience, posted a paean to Modern Language Models (MLMs) entitled Modern language models refute Chomsky’s approach to language on LingBuzz. This triggered a wave of responses from linguists, including one from myself, pointing out the many ways that he was wrong. Recently, Prof. Piantadosi attached a postscript to his paper in which he responds to his critics. The responses are so shockingly bad, I felt I had to respond—at least to those that stem from my critiques—which I will do, spaced out across a few short posts.

In my critique, I brought up the problem of impossible languages, as did Moro et al. in their response. In addressing this critique, Prof. Piantadosi surprisingly begins with a brief diatribe against “poverty of the stimulus.” I say surprisingly, not because it’s surprising for an empiricist to mockingly invoke “poverty of stimulus” much in the same way as creationists mockingly ask why there are still apes if we evolved from them, but because poverty of stimulus is completely irrelevant to the problem of impossible languages and neither I nor Moro et al. even use the phrase “poverty of stimulus.”[1]For my part, I didn’t mention it because empiricists are generally quite assiduous in their refusal to understand poverty of stimulus arguments.

This irrelevancy expressed, Prof. Piantadosi moves on to a more on-point discussion. He argues that it would be wrong-headed for the constraints that would make some languages impossible to be encoded in our model from the start. Rather, if we start with an unconstrained model, we can discover the constraints naturally:

If you try to take constraints into account too early, you might have a harder time discovering the key pieces and dynamics, and could create a worse overall solution. For language specifically, what needs to be built in innately to explain the typology will interact in rich and complex ways with what can be learned, and what other pressures (e.g. communicative, social) shape the form of language. If we see a pattern and assume it is innate from the start, we may never discover these other forces because we will, mistakenly, think innateness explained everything

p36 (v6)

This makes a certain intuitive sense. The problem is that it’s refuted both by the history of generative syntax and the history of science more broadly.

In early theories, a constraint like “No mirroring transformations!” would have to be stated explicitly. Current theories, though, are much simpler with most constraints being derivable from the theory rather than tacked onto the theory.

A digression on scholarly responsibility: Your average engineer working on MLMs could be forgiven for not being up on the latest theories in generative syntax, but Piantadosi is an Associate Professor who has chosen to write a critique of generative syntax, so he really ought to know these things. In fact, he would only not know these thing by a conscious choice not to know or laziness.

Furthermore, the natural sciences have progressed thus far in precisely the opposite direction as what Piantadosi prescribes—they have started with highly constrained theories and progress has generally occurred when some constraint is questioned. Copernicus questioned the constraint that Earth stood still, Newton questioned the constraint that all action was local, Friedrich Wöhler questioned the constraint that organic and inorganic substances were inherently distinct.

None of this, of course, means that we couldn’t do science in the way that Piantadosi suggests—I think Feyerabend was correct that there is no singular Scientific Method—but the proof of the pudding is in the eating. Piantadosi is effectively making a promise that if we let MLM research run its course we will find new insights[2]He seems to contradict himself later on when he asserts that the “science” of MLMs may never be intelligible to humans. More on this in a later post. that we could not find had we stuck with the old direction of scientific progress, and he may be right—just as AGI may actually be 5 years away this time—but I’ll believe it when I see it.


After expressing his methodological objections to considering impossible languages, Piantdosi expresses skepticism as to the existence of impossible languages, stating ” More troubling, the idea of “impossible languages” has never actually been empirically justified.” (p37, v6) This is a truly astounding assertion on his part considering both Moro et al. and I explicitly cite experimental studies that arguable provide exactly the empirical justification that Piantadosi claims does not exist. Both studies cited present participants with two types of made-up languages—one which follows and one which violates the rules of language as theorized by generative syntax—and observes their responses as they try to learn the rules of the particular languages. The study I cite (Smith and Tsimpli 1995) compares the behavioural responses of a linguistic savant to those of neurotypical participants, while the studies cited by Moro et al. (Tettamanti et al., 2002; Musso et al., 2003) uses neuro-imaging techniques. Instead Prof. Piantadosi refers to every empiricists favourite straw-man argument—the alleged lack of embedding structures in Pirahã.

This bears repeating. Both Moro et al. and I expressly point to experimental evidence of impossible languages, and Piantadosi’s response is that no one has ever provided evidence of impossible languages.

So, either Prof. Piantadosi commented on mine and Moro et al‘s critiques without reading them, or he read them and deliberately misrepresented them. It is difficult to see how this could be the result of laziness or even willful ignorance rather than dishonesty.

I’ll leave off here, and return to some of Prof. Piantadosi’s responses to my critiques at a later time.

Notes

Notes
↑1 For my part, I didn’t mention it because empiricists are generally quite assiduous in their refusal to understand poverty of stimulus arguments.
↑2 He seems to contradict himself later on when he asserts that the “science” of MLMs may never be intelligible to humans. More on this in a later post.

The Descriptivist Fallacy

A recent hobby-horse of mine—borrowed from Norbert Hornstein—is the idea that the vast majority of what is called “theoretical generative syntax” is not theoretical, but descriptive. The usual response when I assert this seems to be bafflement, but I recently got a different response—one that I wasn’t able to respond to in the moment, so I’m using this post to sort out my thoughts.

The context of this response was that I had hyperbolically expressed anger at the title of one of the special sessions at the upcoming NELS conference—”Experimental Methods In Theoretical Linguistics.” My anger—more accurately described as irritation—was that, since experiment and theory are complementary terms in science, the title of the session was contradictory unless the NELS organizers were misusing the terms. My point, of course, was that the organizers of NELS—one of the most prestigious conferences in the field of generative linguistics—were misusing the terms because the field as a whole has taken to misusing the terms. A colleague, however, objected, saying that generative linguists were a speech community and that it was impossible for a speech community to systematically misuse words of its own language. My colleague was, in effect, accusing me of the worst offense in linguistics—prescriptivism.

This was a jarring rebuttal because, on the one hand, they aren’t wrong, I was being prescriptive. But, on the other hand and contrary to the first thing students are taught about linguistics, a prescriptive approach to language is not always bad. To see this, let’s consider the to basic rationales for descriptivism as an ethos.

The first rationale is purely practical—if we linguists want to understand the facts of language, we must approach them as they are, not as we think they should be. This is nothing more than standard scientific practice.

The second rationale is a moral one, stemming from the observation that language prescription tends to be directed at groups that lack power in society—Black English has historically been treated as “broken”, features of young women’s speech (“up-talk” in the 90s and “vocal fry” in the 2010s) is always policed, rural dialects are mocked. Thus, prescriptivism is seen as a type of oppressive action. Many linguists make it no further in thinking about prescriptivism, unfortunately, but there are many cases in which prescriptivism is not oppressive. Some good instances of prescriptivism—assuming they are done in good faith—are as follows:

  1. criticizing the use of obfuscatory phrases like “officer-involved shooting” by mainstream media
  2. calling out racist and antisemitic dog-whistling by political actors.
  3. discouraging the use of slurs
  4. encouraging inclusive language
  5. recommending that a writer avoid ambiguity
  6. Asking an actor to speak up

Examples 1 and 2 are obviously non-oppressive uses of prescriptivism, as they are directed at powerful actors; 3 and 4 can be acceptable even if not directed at a powerful person, because they attempt to address another oppressive act; and 5 and 6 are useful prescriptions, as they help the addressee to perform their task at hand more effectively.

Now, I’m not going to try to convince you that the field of generative syntax is some powerful institution, nor that the definition of “theory” is an issue of social justice. Here my colleague was correct—members of the field are free to use their terminology as they see fit. My prescription is of the third variety—a helpful suggestion from a member of the field that wants it to advance. So, while my prescription may be wrong, I’m not wrong to offer it.

Using anti-prescriptivism as a defense against critique is not surprising—I’m sure I’ve had that reaction to editorial suggestions on my work. In fact, I’d say it’s a species of a phenomenon common among folks who care about social justice, where folks mistake a formal transgression for a violation of an underlying principle. In this case the formal act of prescription occurred but without any violation of the principle of anti-oppression.

How do we get good at using language?

Or: What the hell is a figure of speech anyway?

At a certain level I have the same level of English competence as Katie Crutchfield, Josh Gondelman, and Alexandria Ocasio-Cortez. This may seem boastful to a delusional degree of me, but we’re all native speakers of a North American variety of English of a similar age, and this is the level of competence that linguists tend to care about. Indeed, according to our best theories of language, the four of us are practically indistinguishable.

Of course, outside of providing grammaticality judgements, I wouldn’t place myself anywhere near those three, each of whom could easily be counted among the most skilled users of English living. But what does it mean for people to have varied levels of skill in their language use? And is this even something that linguistic theory should be concerned about?

Linguists, of course, have settled on 5 broad levels of description of a given language

  1. Phonetics
  2. Phonology
  3. Morphology
  4. Syntax
  5. Semantics

It seems quite reasonable to say we can break down language skill along these lines. So, skilled speakers can achieve a desire effect by manipulating their phonetics, say by raising their voices, hitting certain sounds in a particular way, or the like. Likewise, phonological theory can provide decent analyses of rhyme, alliteration, rhythm etc. Skilled users of a language also know when to use (morphologically) simple vs complex words, and which word best conveys the meaning they intend. Maybe a phonetician, phonologist, morphologist, or semanticist, will disagree, but these seem like fairly straightforward to formalize, because they all involve choosing from among a finite set of possibilities—a language only has so many lexical entries to choose from. What does skill mean in the infinite realm of syntax? What does it mean to choose the correct figure of speech? Or even more basically, how does one express any figure of speech in the terms of syntactic theory?

It’s not immediately obvious that there is any way to answer these questions in a generative theory for the simple reason that figures of speech are global properties of expressions, while grammatical theory deals in local interactions between parts of expressions. Take an example from Abraham Lincoln’s second inaugural address:

(1) Fondly do we hope—fervently do we pray—that this mighty scourge of war may speedily pass away.

There are three syntactic processes employed by Lincoln here that I can point out:

(2) Right Node Raising
Fondly do we hope that this mighty scourge of war may speedily pass away, and fervently do we pray that this mighty scourge of war may speedily pass away. -> (1)

(3) Subject-Aux Inversion
Fondly we hope … -> (1)

(4) Adverb fronting
We hope fondly… -> (1)

Each of these represents a choice—conscious or otherwise—that Lincoln made in writing his speech and, while most generative theories allow for choices to be made, they are not at the same levels.

Minimalist theories, for instance, allow for choices at each stage of sentence construction—you can either move constituent, add a constituent, or stop the derivation. Each of (3) and (4) could conceivably be represented as a single choice, but it seems highly unlikely that (2) could. In fact, there is nothing approaching a consensus as to how right node raising is achievable, but it is almost certainly a complex phenomenon. It’s not as if we have a singular operation RNR(X) which changes a mundane sentence into something like (1), yet Lincoln and other writers and orators seem to have it as a tool in their rhetorical toolboxes.

Rhetorical skill of this kind suggest the possibility of a meta-grammatical knowledge, which all speakers of a language have to some extent, and which highly skilled users have in abundance. But what could this meta-grammatical knowledge consist of? Well, if the theoretical representation of a sentence is a derivation, then the theoretical representation of a figure of speech would be a class of derivations. This suggests an ability to abstract over derivations in some way and therefore, it suggests that we are able to acquire not just lexical items, but also abstractions of derivations.

This may seem to contradict the basic idea of Minimalism by suggesting two grammatical systems and indeed, it might be a good career move on my part to declare that the fact of figures of speech disproves the SMT, but I don’t see any contradiction inherent here. In fact, what I’m suggesting here and have argued for elsewhere is something that is a fairly basic observation from computer science and mathematical logic—that the distinction between operations and operands is not that distinct. I am merely suggesting that part of a mature linguistic knowledge is higher-order grammatical functions—functions that operate on other functions and/or yield other functions—and that, since any recursive system is probably able to represent higher-order functions, we should absolutely expect our grammars to allow for them.

Assuming this sort of abstraction is available and responsible for figures of speech, our task as theorists then is to figure out what form the abstraction takes, and how it is acquired, so I can stop comparing myself to Katie Crutchfield, Josh Gondelman, and AOC.

De re/De dicto ambiguities and the class struggle

If you follow the news in Ontario, you likely heard that our education workers are demanding an 11.7% wage raise in the current round of bargaining with the provincial government. If, however, you are more actively engaged with this particular story—i.e., you read past the headline, or you read the union’s summary of bargaining proposals—you may have discovered that, actually, the education workers are demanding a flat annual $3.25/hr increase across the board. On the surface, these seem to be two wildly different assertions that can’t both be true. One side must be lying! Strictly speaking, though, neither side is lying, but one side is definitely misinforming.

Consider a version of the headline (1) that supports the government’s line.

(1) Union wants 11.7% raise for Ontario education workers in bargaining proposal.

This sentence is ambiguous. More specifically is shows a de re/de dicto ambiguity. The classic example of such an ambiguity is in (2).

(2) Alex wants to marry a millionaire.

There is one way of interpreting this in which Alex wants to get married and one of his criteria for a spouse is that they be a millionaire. This is the de dicto (lit. “of what is said”) interpretation of (2). The other way of interpreting it is that Alex is deeply in love with a particular person and wants to marry them. It just so happens that Alex’s prospective spouse is a millionaire—a fact which Alex may or may not know. This is the de re (lit. “of the thing”) interpretation of (2). Notice how (2) can describe wildly different realities—for instance, Alex can despise millionaires as a class, but unknowingly want to marry a millionaire.

Turning back to our headline in (1), what are the different readings? The de dicto interpretation is one in which the union representatives sit down at the bargaining table and say something like “We demand an 11.7% raise”. The de re interpretation is one in which the union representatives demanded, say, a flat raise that happens to come out to an 11.7% raise for those workers with the lowest wages when you do the math. The de re interpretation is compatible with the assertions made by the union, so it’s probably the accurate interpretation.

So, (1) is, strictly speaking, not false under one interpretation. It is misinformation, though, because it deliberately introduces a substantive ambiguity in a way that, the alternative headline in (3) does not.

(3) Union wants $3.25/hr raise for Ontario education workers in bargaining proposal

Of course (3) has the de re/de dicto ambiguity—all expressions of desire do—but both interpretations would accurately describe the actual situation. Someone reading the headline (3) would be properly informed regardless of how they interpreted it, while (1) leads some readers to believe a falsehood.

What’s more, I think it’s reasonable to call the headline in (1) deliberate misinformation.

The simplest way to report the union’s bargaining positions would be to simply report it—copy and paste from their official summary. To report the percentage increase as they did, someone had to do the arithmetic to convert absolute terms to relative terms—a simple step, but an extra step nonetheless. Furthermore, to report a single percentage increase, they had to look only at one segment of education workers—the lowest-paid segment. Had they done the calculation on all education workers, they would have come up with a range of percentages, because $3.25 is 11.7% of $27.78, but 8.78% of 37.78, and so on. So, misinforming the public by publishing (1) instead of (3) involved at least two deliberate choices.

It’s worth asking why misinform in this way. A $3.25/hr raise is still substantial and the government could still argue that it’s too high, so why misinform? One reason is that puts workers in the position of explaining that it’s not a bald-faced lie, but it’s misleading, making us seem like pedants. but I think there’s another reason for the government to push the 11.7% figure, it plays into and furthers an anti-union trope that we’re all familiar with.

Bosses always paint organized labour as lazy, greedy, and corrupt—”Union leaders only care about themselves only we bosses care about workers and children.” They especially like to claim that unionized workers, since they enjoy higher wages and better working conditions, don’t care about poor working folks.[1]Indeed there are case in which some union bosses have pursued gains for themselves at the expense of other workers—e.g., construction Unions endorsing the intensely anti-worker Ontario PC Party … Continue reading The $3.25/hr raise demand, however, reveals these tropes as lies.

For various reasons, different jobs, even within a single union, have unequal wages. These inequalities can be used as a wedge to keep workers fighting amongst themselves rather than together against their bosses. Proportional wage increases maintain and entrench those inequalities—if everyone gets a 5% bump, the gap between the top and bottom stays effectively the same. Absolute wage increases, however, shrink those inequalities. Taking the example from above a $37.78/hr worker makes 1.33x the $27.78/hr worker, but after a $3.25/hr raise for both the gap narrow slightly to 1.29x, and continues to do so. So, contrary to the common trope, union actions show solidarity rather than greed.[2]Similar remarks can be made about job actions, which are often taken as proof that workers are inherently lazy. On the contrary, strikes are physically and emotionally grueling and rarely taken on … Continue reading

So what’s the takeaway here? It’s frankly unreasonable to expect ordinary readers to do a formal semantic analysis of their news, though journalists could stand to be a bit less credulous of claims like (1). My takeaway is that this is just more evidence of my personal maxim that people in positions of power lie and mislead whenever it suits them as long as no one questions them. Also, maybe J-schools should have required Linguistics training.

Notes

Notes
↑1 Indeed there are case in which some union bosses have pursued gains for themselves at the expense of other workers—e.g., construction Unions endorsing the intensely anti-worker Ontario PC Party because they love building pointless highways and sprawling suburbs
↑2 Similar remarks can be made about job actions, which are often taken as proof that workers are inherently lazy. On the contrary, strikes are physically and emotionally grueling and rarely taken on lightly