A Statistical Analysis of the narrators of ‘Ulysses’ or ‘why ‘Ulysses’ isn’t wisdom literature’

The second time I read Ulysses,in advance of an undergraduate seminar, it was around the ninetieth anniversary of the original text’s publication. The newspapers were printing archive material relating to the novel, extended supplements about its importance from the usual quarters, as well as reviews of recently published monographs from both young and established scholars. Unfortunately, the critical trend of the time was to read Ulysses as wisdom literature. Critics urged prospective readers of the novel to wrest Joyce from the scholars and bring him ‘back to the people’. This school of thought treated Leopold Bloom as a model of the way in which the contemporary urban subject should be living: aloof, polite, well-intentioned but not dogmatic on political issues. Moderately informed, but more often wrong, a reader, but not self-serious, an everyman. Ulysses’ structural indebtedness to cornerstones of The Canon such as William Shakespeare’s Hamlet and Homer’s The Odyssey frequently undergirds this line of argument, demonstrative in itself of how easily high literary art and everyday life may be set next to one another. This generally requires critics to treat the characters of Bloom and Stephen Dedalus as two opposites in need of the other. Each has a little to impart on life, love and literature, whether it be to reflect a little deeper on themselves or their marriage, move past their respective losses or to find in each other their lost son/father.

This interpretation of the novel reads it along a linear trajectory, as Stephen and Bloom come together to form Blephen and Stoom. Through computation it may be possible to examine the writing style of later chapters, and determine whether or not they bear formal witness to this change in character. We must first however, consider the difficulty of locating where Joyce’s narrators actually are. Part of what makes Joyce’s writing style so unique is his use of free indirect discourse, a mode of writing in which the reality of the text is inflected by the consciousness(es) of the beholder(s). As such, putting a category on each episode of Ulysses as though it were narrated by one person or a combination of persons might seem reductive; it very much is. But in fusing computation and literature, certain assumptions have to be made.

In carrying out this analysis, I made use of R’s ‘Stylo’ package, which contains tools for breaking a number of texts into equal sizes, removing words which are not common to most samples, calculating the relative frequencies of these words, transforming these observations into new combinations of variables called ‘components’ with greater explanatory potential, and clustering them together. These words appear below:

These might seem like boring terms, as literary critics we tend to look past them to more evocative ones like ‘serpentine’ or ‘columbanus’ but unfortunately, in computational terms it is the relative frequencies of these ‘particles’ or ‘function words’ that provide the most secure means of modelling a writer’s particular idiom. These samples were then plotted on a correlation matrix, which can be taken as an index of similarity, based on where they cluster:

The six different narrators of Ulysses appearing in the index above are:

‘Anon’, who narrates the episode ‘Cyclops’

‘Blephen’, a composite delineation for episodes in which both characters feature, such as ‘Circe’, ‘Eumaeus’, ‘Ithaca’ and ‘Oxen of the Sun’

Bloom, who narrates ‘Hades’, ‘Calypso’, ‘Lestrygonians’ and ‘The Lotus Eaters’, Gerty, who narrates at least half of ‘Nausicaa’ (this is a controversial point within the literature, it might by Bloom who is narrating for her)

Molly, who narrates the book’s final chapter ‘Penelope’,

and finally Stephen, who narrates the first three episodes ‘Telemachus’, ‘Nestor’ and ‘Proteus’, as well as the novel A Portrait of the Artist as a Young Man, which has been thrown in here for comparison.

Here’s the same plot as above with the labels more clearly indicated

The first thing we could note is the gender divide. Molly and Gerty both spread over to the right, with Molly as an outlier. Both are more proximate to the A Portrait samples than any other, which are all taken from the earlier parts of the novel, suggesting that Joyce writes women and young children using the same number of words at the same rate. As the Gerty samples move through the episode, they move closer and closer to the Bloom cluster, visually conforming that the episode starts in Gerty’s voice before he takes over, and that Bloom doesn’t think much of women’s intelligence in the main either.

Overall we can say that there doesn’t look to be a fusing of perspectives here as such. Rather than the Blephen episodes meeting halfway between the Stephen and Bloom, Stephen and Bloom already seem quite comfortably clustered at the novel’s outset. Based on the divide between Stephen’s episodes of Ulysses and A Portrait, we might say that the way in which Stephen narrates A Portrait is very different from the way in which he narrates Ulysses.This is justified I think by how sensitive the analysis is to changes in narrator, demonstrated by the Gerty/Bloom example already discussed, as well as the fact that the earlier part of Aeolous, in which Bloom is present, clusters with his samples, whereas the second part, after Stephen’s entered, clusters with the Stephen samples.

Below is the plot with the Portrait samples removed:

Words Stephen’s narration is most likely to use in comparison to Bloom
Words Bloom’s narration is more likely to use in comparison to Stephen

There are a number of ways one could use these results to interrogate the notion of Ulysses as wisdom literature. We could begin by asking after the gendered aspects of the adjective ‘wise’, and ask why so many of these books which teach us how one might best live are written by men (and how tone-deaf this argument can sound because to read Ulysses one might almost think married women weren’t let out of the house) or we could ask what interests an Irish model of bourgeois respectability might serve, along the lines of an Irish ‘keep calm and carry on’ poster.

Ulysses as a guide to life risks rendering it a novel of parts coming together, the middle-class intellectual and the middle-class working stiff holding hands across whatever barricade is supposed to be dividing them. Not that I would go to the other extreme and frame it as one of dissolution. Ulysses’ shape is one I would be loathe to put a vector to in fact; to say that Stephen and Bloom’s relationship moves from a) state to b) state would be too easy by half.

What makes Ulyssesan interesting novel to me is its self-referentiality, the dialogue it establishes between the novel and its supposed referent of ‘real Dublin’, which is made most clear in ‘Circe’, but also in the book’s other failed attempts to understand itself, as in the cases of the characters referenced as being in particular places at particular times who may or may not be Bloom, the McIntosh mystery or the puzzle of crossing Dublin without passing a pub. In this context, I think ‘Eumaeus’ appearing as a stylistic outlier is significant.

It is in this episode that we get information about a sequence of coincidences, and resonant differences between Bloom and Stephen’s lives. The depth of these coincidences (which I won’t provide a summary of here, because I think they’re among the most poignant parts of the novel) gesture towards something a bit more cosmically ordered than the rest of the novel even as they take place within the circumscribed rituals of Irish urban middle-class life in the early twentieth century. ‘Eumaeus’ is written in a chill tone which most closely resembles that of a scientific paper, eliding the indirect discourse which ostensibly defines the rest of the text, and it is the fact that these connections are raised here rather than anywhere else that the true interest in their relationship, such as it is, is to be found.

These connections which remain unrealised by the two, rather than bring us to some Forsterian notion of connection should raise instead questions of alienation and of their unity in separation. It presents problems both epistemological and political, about how our reality is structured, the means through which it is circumscribed and how it is more defined by how little of it we are aware of rather than how much. Rather than teaching us ‘how to live’ Ulysses shows us how we do not live, how we probably won’t live and how it could so easily have been otherwise. It is no more an explanation for life as it is an explanation of itself, or Homer, or Ireland.


Quantifying Modernism and the avant-garde

Introduction and Methodology

(Skip to results if you want to miss the boring parts, or look here for a more granular, in depth account, including the code itself. If you code, yeah, I’m so sorry, I’ll make it more elegant soon)

This post will document a statistical analysis which was carried out on a corpus of 500 novels. 250 of these texts are generally categorised as ‘realist’ and will be used as a benchmark against which we might define modernist literary style, a mode of writing which arose in the early twentieth century, (though it should be noted that this chronology is increasingly subject to revision due to the work of new modernist scholars).

The first novel in the naturalistic corpus, chronologically speaking, is Jane Austen’s novel Lady Susan, and was written in the year 1794. The final one is Thomas Hardy’s novel Jude the Obscure, which was published in 1895. This corpus contains the complete prose works, a phrase here encompassing novels, novellas and short story collections, of fifteen writers, Jane Austen, Emily, Anne and Charlotte Bronte, Stephen Crane, Honoré de Balzac, Charles Dickens, Fyodor Dostoevsky, George Eliot, Gustave Flaubert, Elizabeth Gaskell, Thomas Hardy, William Makepeace Thackeray, Leo Tolstoy and Émile Zola.

The corpus of 250 modernist novels begins in the year 1869, with Henry James’ first bloc of short stories, and continues all the way to Samuel Beckett’s 1988 novella ‘Stirrings Still’, so there is some overlap between these two corpora’s starting and end points. This modernist corpus otherwise consists of the complete works of nineteen writers such as Djuna Barnes, Samuel Beckett, Jorge Luis Borges, Elizabeth Bowen, Joseph Conrad, William Faulkner, F. Scott FitzGerald, Ford Madox Ford, Ernest Hemingway, Henry James, James Joyce, Franz Kakfa, D.H. Lawrence, Katherine Mansfield, Flann O’Brien, Marcel Proust, Gertrude Stein, Edith Wharton and Virginia Woolf.

This disproportion between the two corpora, with fifteen realists versus ninteen modernists, may seem disconcerting at first, but what is required in order for the statistical analyses to function is for the number of observations to be equal, rather than the number of novelists. Unfortunately, realist authors wrote more novels than modernist authors, and this compromised our ability to retain the same number of authors on each end of the generic spectrum.

One other aspect to consider is the international dimension. The realist corpus includes ten novelists who wrote in English, but there are also two Russian and three French realists, two of whom, Zola and the aforementioned Balzac, were far more prolific than any other writer in either corpus. Zola and Balzac composed 86 and 34 novels, short story collections or novellas respectively. This has the consequence that well over half of the realist corpus is in translation from another language in comparison to just under 10% of the modernist corpus. I intend to address this when I am at a later stage in my research. There has been some work published on the issues surrounding the quantification of literature in translation and across language, but I do not yet possess a sufficient breadth of knowledge in this field to comment intelligently on the matter. I do think it is important to have French and Russian writers included in the realist corpus on the basis that many of them, be they Tolstoy, Flaubert or Balzac, exerted a significant influence on their modernist successors.

Whether or not these are ‘the best’ or most accurate translations is sort of beside the point, from the reading I have done around the issue of literary translation, their being subject to change over time is in the nature of how text is received and re-constituted in different eras for different communities of readers (this discussion between Will Self and Kafka’s translators is particularly illuminating in this context, please do not be put off by Self, he gives the translators so much space to discuss the process, you really should watch it). The germane point here is that the translations being analysed in this instance could not be considered to be the most contemporary. There might be an argument for retaining these older translations on the basis that they are more likely to be the versions of the text which would have been circulating in the early twentieth century and therefore the translations modernist authors would have been more likely to have read, but making this claim would require a greater burden of proof, such as what languages each author read novels in and what their reading habits were more generally.

So, to turn to the analysis. My research is directed towards the quantitative analysis of grammar, the rationale being that we could, by examining varying quantities of particular categories of words, such as verbs, adjectives or prepositions, develop an understanding of how literary fiction changes from the beginning of the nineteenth century until the end of the twentieth, and, more specifically, how literary modernism departs from, or, perhaps remains contiguous with, this previous generation of novel writing. This was carried out using a POS tagger from the Natural Language Toolkit in Python.


From realism to modernism:

  • average sentence length decreases by 4 words, from an average 22 words to 18 words per sentence.
  • Personal pronouns (I, you, he, she, it, we, they, me, him, her, us, and them) increase by 1% from 5% to 6%. Interrogative pronouns (who and where) also decrease by 0.01% from 0.03% to 0.02%
  • Verbs in the past tense increase by 1% from 6% to 7%.
  • Adverbs increase by 0.5% from 4.5% to 5%.
  • Prepositions, (after, in, to, on, and with) decrease by 0.4% from 10.9% to 10.5%
  • Wh Determiners (words beginning with wh, such as ‘where’ or ‘who’ acting to modify the noun phrase) decrease by 0.2% from 0.6% to 0.4%.
  • Particles (parts of speech with grammatical function with no meaning such as ‘up’ in the phrase ‘I tidied up the room’) increase by 0.1% from 0.4% to 0.5%.
  • Non third-person singular present verbs (verbs in first or second person) decrease by 0.1% from 1.6% to 1.5%.
  • Existentials (words such as ‘there’ which indicates that something exists) increase by 0.04%, from 0.17% to 0.21%.
  • Superlative adjectives (adjectives such as ‘best’, ‘biggest’, ‘worst’) decrease by 0.01% from 0.14% to 0.13%.

It will not have escaped your attention that a lot of these percentages are quite small. The extent to which any given text is made up of this hyper-specific categories is pretty minimal in the first place, so this is why many of these quantities seem so laughably tiny. Rest assured that they are statistically significant, this does not mean that they are important, this requires a greater burden of proof, more analyses, more exploration, but that they are noteworthy considering the quantities involved.

One boxplot which might be of interest, is the one below, which shows the ‘spread’ of the data for average sentence length between realism and modernism.

What we see on the left is the variation of the sentence length data (the term ‘variation’ here meaning the general ‘dispersedness’ of the data) for realism, which goes from 10 to roughly 35 words per sentence with an outlier or two on either end, whereas if we consider modernism, we have everything from zero (Samuel Beckett’ novel How It Is which has no full stops in it) up to forty-five, with far more outliers on the higher end. Higher outliers, are data points with values greater than 1.5 times the interquartile range above the third quartile, lower outliers, of which there are three, are more than 1.5 times below the first quartile. For one’s own general knowledge, the modernist outliers for sentence length are

  • William Faulkner’s Absalom! Absalom! (46.4), and Intruer in the Dust (42.3)
  • Marcel Proust’s Swann’s Way (42.9), In a Budding Grove (40.2) In a Budding Grove (40.2), Time Re-gained (38), The Prisoner (37.2) and The Captive (35.7) The Guermantes Way (34.1) and Sodom and Gomorrah (30.9).
  • Samuel Beckett’s Texts for Nothing and The Unnamable have 40.5 and 32.9 words per sentence respectively
  • Gertrude Stein’s novels The Making of Americans and Everybody’s Autobiography have 33.9 and 33.5 respectively.
  • Henry James’ The Ivory Tower and The Young Lovell score 31.8 and 29 respectively.
  • The three lower outlier values for sentence length are all written by Beckett, such as the aforementioned How It Is and also Worstward Ho (4.9) and Ill Seen Ill Said (7).

It can be tempting I think, when we see these sorts of names surface so prominently, in conjunction with a visual confirmation of the existence of an avant-garde to think that modernism in its most pure form was a kind of relentless maximalism, an uncompromising movement towards longer sentences, more pronouns, and that all other manifestations of it are inadequate or insufficient in some way. This is a kind of a boring and masculinist overview of the genre, which takes, I think, too many of the claims made by its most dogmatic adherents at face value, and it’s not a modernism I’m particularly interesting in defending or instantiating. There can also, of course, be a regressive or rearguard aspect to modernism, which is perceptible in the following boxplot, which displays the distribution of past tense verbs.

As was pointed out above, modernism displays an increase in past tense verbs overall, but here we see a large number of outlier values moving against the overall trend. These novels are:

  • James Joyce’s Ulysses (4.3%) and Finnegans Wake (2.7%)
  • William Faulkner’s As I Lay Dying (4.2%) and Requiem for a Nun (3.6%)
  • Samuel Beckett’s Malone Dies (3.9%), Fizzles (2.5%), Company (2%), Texts for Nothing (1.8%), The Unnamable (1.7%), Worstward Ho (1.6%), Ill Seen Ill Said (1.4%) and a corpus of his miscellaneous and unpublished short fiction (2.2%).
  • Joseph Conrad and Ford Madox Ford’s collaborative novel The Nature of a Crime (2.6%)
  • Virginia Woolf’s The Waves (2.4%)
  • Gertrude Stein’s Tender Buttons (1.7%)

The higher modernism outlier is Virginia Woolf’s 1937 novel The Years (10%) and the lower realism outlier is Balzac’s 1841 novel Letters of Two Brides (2.7%)

In this way we can see that modernism is not just a unidirectional commitment to a narrow sequence of stylistic changes. Instead, it’s a contradictory movement in which a number of different stylistic markers jostle against and subvert one another. In this particular instance, for example, we can perceive the authors most generally understood to be among the most uncompromising; Joyce, Beckett, Stein, Woolf and Faulkner, resisting the overall trend.

From the two boxplots I’ve generated so far, you might have noticed that in, modernism tends to generate a greater number of outliers, and I can confirm that this trend of a greater degree grammatical heterogeneity manifesting itself in modernist novel-writing than naturalistic novel-writing persists across the other categories of grammar, which you can validate by looking at the complete analysis here.

This struck me as important development, so I quantified the extent of each data point’s outlier-ness, and then grouped them according to author. These values were then divided by the number of outlier data points, because some of these novelists only have a small number of novels in the corpus versus others. Austen’s complete works would be totally outnumbered by Balzac’s for instance. The results appear below:

Please do note the values on the y-axis; Jane Austen is barely above zero because the only outlier text she wrote is Mansfield Park, which marks itself out for its disproportional use of adjectives. I thought it better to not exclude her from the plot though, because, I didn’t want it to turn into even more of a boy’s club than it might otherwise be. It would be useful, and exciting I think, to conceive of this plot as an indication of early breaches with conventional form, perhaps some nineteenth century anticipations of modernism. Reading Dostoevsky, Zola and Balzac in this manner would all be coterminous with changes taking place in the study of modernism now, but reading Thackeray and Eliot in these terms might be a more surprising development, and I’d be interested to read these texts in light of what we’re seeing here.

The modernism plot for deviation appears below:

The unlabelled entry between Faulkner and James is Hemingway

From this plot we can see that the most avant-gardist prose writers, considered from the perspective of their grammar, appear to be Beckett, Stein, Woolf, Conrad and Joyce. Of course, this is nowhere near a definitive answer as to what modernist style is, or who its most innovative practitioners were; these measurements are atomistic and are quantifying individual words. But style is not just words in isolation, style is agglomerations of words, spaces between words, the clandestine networks and relations the phrases these words add up to compose in the mind of the reader, and, if these digital methodologies are to have any chance of illustrating this shift (an inadequate term in the first instance, since it is more an accumulation of changes distributed over a broad corpus than a sudden or transformational one that we are here concerned with) it is in these cumulative terms that style must be quantified, in order to avoid drifting into the reductive and schematic scientism that numerical analyses of this kind are frequently accused of perpetuating.

How big are the words modernists use?

It’s a fairly straightforward question to ask, one which most literary scholars would be able to provide a halfway decent answer to based on their own readings. Ernest Hemingway, Samuel Beckett and Gertrude Stein more likely to use short words, James Joyce, Marcel Proust and Virginia Woolf using longer ones, the rest falling somewhere between the two extremes.

Most Natural Language Processing textbooks or introductions to quantitative literary analysis demonstrate how the most frequently occurring words in a corpus will decline at a rate of about 50%, i.e. the most frequently occurring term will appear twice as often as the second, which is twice as frequent as the third, and so on and so on. I was curious to see whether another process was at work for word lengths, and whether we can see a similar decline at work in modernist novels, or whether more ‘experimental’ authors visibly buck the trend. With some fairly elementary analysis in NLTK, and data frames over into R, I generated a visualisation which looked nothing like this one.*

*The previous graph had twice as many authors and was far too noisy, with not enough distinction between the colours to make it anything other than a headwreck to read.

In narrowing down the amount of authors I was going to plot, I did incline myself more towards authors that I thought would be more variegated, getting rid of the ‘strong centre’ of modernist writing, not quite as prosodically charged as Marcel Proust, but not as brutalist as Stein either. I also put in a couple of contemporary writers for comparison, such as Will Self and Eimear McBride.

As we can see, after the rather disconnected percentages of corpora that use one letter words, with McBride and Hemingway on top at around 25%, and Stein a massive outlier at 11%, things become increasingly harmonious, and the longer the words get, the more the lines of the vectors coalesce.

Self and Hemingway dip rather egregiously with regard to their use of two-letter words (which is almost definitely because of a mutual disregard for a particular word, I’m almost sure of it), but it is Stein who exponentially increases her usage of two and three letter words. As my previous analyses have found, Stein is an absolute outlier in every analysis.

By the time the words are ten letters long, true to form it’s Self who’s writing is the only one above 1%.

Collocations in Modernist Prose

Screen Shot 2017-07-24 at 14.51.47I have recently begun to experiment with Natural Language Processing to determine how particular words in modernist texts are correlated. I’m still getting my head around Python and NLTK, but so far I’m finding it much more user-friendly than similar packages in R.

Long-term I hope to graph these collocations in high-vector space, so that I can graph them, but for the moment, I’m interested in noting the prevalence of the term ‘young man’, Self and Baume being the only authors that have female adjective-noun phrases, and the usage of titles which convey particular social hierarchies; Joyce, Woolf and Bowen’s collocations are almost exclusively composed of these, as is Stein’s, with the clarifier that Stein’s appear shorn of their ‘Mr.’, ‘Miss.’ or ‘Doctor’.

Here’s all the collocations in the modernist corpus:

young man; robert jordan; new york; gertrude stein; old man; could see; henry martin; every one; years ago; first time; long time; hugh monckton; great deal; come back; david hersland; good deal; every day; edward colman; came back; alfred hersland

Canonical modernist texts:

young man; robert jordan; gertrude stein; henry martin; new york; every one; old man; could see; years ago; long time; hugh monckton; first time; great deal; david hersland; come back; good deal; every day; edward colman; alfred hersland; mr. bettesworth

Contemporary texts, Enright, Self, Baume, McBride:

fat controller; phar lap; von sasser; first time; per cent; could see; old man; one another; even though; years ago; new york; front door; young man; either side; someone else; dave rudman; last night; living room; steering wheel; every time

Djuna Barnes

frau mann; nora said; english girl; someone else; long ago; leaned forward; london bridge; come upon; could never; god knows; doctor said; sweet sake; first time; five francs; terrible thing; francis joseph; hôtel récamier; orange blossoms; bowed slightly; would say

Eimear McBride

kentish town; someone else; first time; last night; jesus christ; something else; years ago; five minutes; every day; hail mary; take care; next week; arms around; never mind; every single; little girl; little boy; two years; soon enough; come back

Elizabeth Bowen

mrs kerr; lady waters; mrs heccomb; major brutt; mme fisher; lady naylor; miss fisher; good deal; said mrs; first time; lady elfrida; one another; young man; colonel duperrier; aunt violet; last night; ann lee; one thing; sir robert; sir richard

Ernest Hemingway

robert jordan; old man; could see; colonel said; gran maestro; catherine said; jordan said; richard gordon; long time; pilar said; thou art; pablo said; nick said; bill said; girl said; captain willie; young man; automatic rifle; mr. frazer; david said

F. Scott FitzGerald

new york; young man; years ago; first time; sally carrol; several times; fifth avenue; ten minutes; minutes later; richard caramel; thousand dollars; five minutes; young men; evening post; old man; next day; saturday evening; long time; last night; come back

Gertrude Stein

gertrude stein; every one; david hersland; alfred hersland; angry feeling; family living; independent dependent; jeff campbell; julia dehning; mrs. hersland; daily living; whole one; bottom nature; madeleine wyman; good deal; mary maxworthing; middle living; miss mathilda; mabel linker; every day

James Joyce

buck mulligan; said mr.; martin cunningham; aunt kate; says joe; mary jane; corny kelleher; ned lambert; mrs. kearney; stephen said; mr. henchy; ignatius gallaher; father conmee; nosey flynn; mr. kernan; myles crawford; cissy caffrey; ben dollard; mr. cunningham; miss douce

Marcel Proust

young man; faubourg saint-germain; long ago; caught sight; first time; every day; one day; great deal; des laumes; young men; could see; quite well; next day; one another; would never; nissim bernard; victor hugo; would say; louis xiv; long time

Samuel Beckett

said camier; said mercier; miss counihan; lord gall; miss carridge; mr. kelly; panting stops; said belacqua; mr. endon; said wylie; said neary; one day; otto olaf; dr. killiecrankie; come back; vast stretch; mrs gorman; push pull; something else; ground floor

Sara Baume

even though; tawny bay; living room; old man; passenger seat; bird walk; maggot nose; shut-up-and-locked room; stone fence; food bowl; lonely peephole; low chair; old woman; kennel keeper; rearview mirror; shih tzu; shore wall; safe space; every day; oneeye oneeye

Virginia Woolf

miss barrett; mrs. ramsay; mrs. hilbery; young man; st. john; could see; years ago; peter walsh; mrs. thornbury; miss allan; said mrs.; young men; mrs. swithin; human beings; wimpole street; mrs. flushing; mr. ramsay; mrs. manresa; sir william; door opened

Anne Enright

new york; per cent; eliza lynch; dear friend; years old; even though; first time; came back; years ago; long time; michael weiss; señor lópez; living room; every time; looked like; could see; one day; said constance; pat madigan; mrs hanratty

Will Self

fat controller; phar lap; von sasser; one another; old man; could see; first time; per cent; dave rudman; let alone; front door; young man; skip tracer; quantity theory; jane bowen; los angeles; young woman; either side; charing cross; long since

Flann O’Brien

father fahrt; good fairy; father cobble; said shanahan; mrs crotty; said furriskey; said lamont; mrs laverty; one thing; sergeant fottrell; said slug; old mathers; public house; far away; cardinal baldini; monsignor cahill; mrs furriskey; red swan; black box; said shorty

Ford Madox Ford

henry martin; hugh monckton; edward colman; privy seal; mr. bettesworth; mr. fleight; young man; mr. sorrell; sergius mihailovitch; young lovell; new york; jeanne becquerel; lady aldington; kerr howe; anne jeal; miss peabody; mr. pett; great deal; marie elizabeth; robert grimshaw

Jorge Luis Borges

ts’ui pên; buenos aires; pierre menard; eleventh volume; richard madden; nils runeberg; yiddische zeitung; stephen albert; hundred years; erik lönnrot; firing squad; henri bachelier; madame henri; orbis tertius; vincent moon; paint shop; seventeenth century; anglo-american cyclopaedia; fergus kilpatrick; years ago

Joseph Conrad

mrs. travers; mrs verloc; mrs. fyne; peter ivanovitch; doña rita; miss haldin; mrs. gould; assistant commissioner; charles gould; san tomé; chief inspector; years ago; captain whalley; could see; van wyk; old man; dr. monygham; gaspar ruiz; young man; mr. jones

D.H. Lawrence

young man; st. mawr; mr. may; mrs. witt; blue eyes; miss frost; could see; one another; mrs bolton; ‘all right; come back; said alvina; two men; of course; good deal; long time; mr. george; next day

William Faulkner

uncle buck; aleck sander; miss reba; years ago; dewey dell; mrs powers; could see; white man; four years; old man; ned said; division commander; general compson; miss habersham; new orleans; uncle buddy; let alone; one another; united states; old general

Literary Cluster Analysis

I: Introduction

My PhD research will involve arguing that there has been a resurgence of modernist aesthetics in the novels of a number of contemporary authors. These authors are Anne Enright, Will Self, Eimear McBride and Sara Baume. All these writers have at various public events and in the course of many interviews, given very different accounts of their specific relation to modernism, and even if the definition of modernism wasn’t totally overdetermined, we could spend the rest of our lives defining the ways in which their writing engages, or does not engage, with the modernist canon. Indeed, if I have my way, this is what I will spend a substantial portion of my life doing.

It is not in the spirit of reaching a methodology of greater objectivity that I propose we analyse these texts through digital methods; having begun my education in statistical and quantitative methodologies in September of last year, I can tell you that these really afford us no *better* a view of any text then just reading them would, but fortunately I intend to do that too.

This cluster dendrogram was generated in R, and owes its existence to Matthew Jockers’ book Text Analysis with R for Students of Literature, from which I developed a substantial portion of the code that creates the output above.

What the code is attentive to, is the words that these authors use the most. When analysing literature qualitatively, we tend to have a magpie sensibility, zoning in on words which produce more effects or stand out in contrast to the literary matter which surrounds it. As such, the ways in which a writer would use the words ‘the’, ‘an’, ‘a’, or ‘this’, tends to pass us by, but they may be far more indicative of a writer’s style, or at least in the way that a computer would be attentive to; sentences that are ‘pretty’ are generally statistically insignificant.

II: Methodology

Every corpus that you can see in the above image was scanned into R, and then run through a code which counted the number of times every word was used in the text. The resulting figure is called the word’s frequency, and was then reduced down to its relative frequency, by dividing the figure by total number of words, and multiplying the result by 100. Every word with a relative frequency above a certain threshold was put into a matrix, and a function was used to cluster each matrix together based on the similarity of the figures they contained, according to a Euclidean metric I don’t fully understand.

The final matrix was 21 X 57, and compared these 21 corpora on the basis of their relative usage of the words ‘a’, ‘all’, ‘an’, ‘and’, ‘are’, ‘as’, ‘at’, ‘be’, ‘but’, ‘by’, ‘for’, ‘from’, ‘had’, ‘have’, ‘he’, ‘her’, ‘him’, ‘his’, ‘I’, ‘if’, ‘in’, ‘is’, ‘it’, ‘like’, ‘me’, ‘my’, ‘no’, ‘not’, ‘now’, ‘of’, ‘on’, ‘one’, ‘or’, ‘out’, ‘said’, ‘she’, ‘so’, ‘that’, ‘the’, ‘them’, ‘then’, ‘there’, ‘they’, ‘this’, ‘to’, ‘up’, ‘was’, ‘we’, ‘were’, ‘what’, ‘when’, ‘which’, ‘with’, ‘would’, and ‘you’.

Anyway, now we can read the dendrogram.

III: Interpretation

Speaking about the dendrogram in broad terms can be difficult for precisely the reason that I indicative above; quantitative/qualitative methodologies for text analysis are totally opposed to one another, but what is obvious is that Eimear McBride and Gertrude Stein are extreme outliers, and comparable only to each other. This is one way unsurprising, because of the brutish, repetitive styles and is in other ways very surprising, because McBride is on record as dismissing her work, for being ‘too navel-gaze-y.’

Jorge Luis Borges and Marcel Proust have branched off in their own direction, as has Sara Baume, which I’m not quite sure what to make of. Franz Kafka, Ernest Hemingway and William Faulkner have formed their own nexus. More comprehensible is the Anne Enright, Katherine Mansfield, D.H. Lawrence, Elizabeth Bowen, F. Scott FitzGerald and Virginia Woolf cluster; one could make, admittedly sweeping judgements about how this could be said to be modernism’s extreme centre, in which the radical experimentalism of its more revanchiste wing was fused rather harmoniously with nineteenth-century social realism, which produced a kind of indirect discourse, at which I think each of these authors excel.

These revanchistes are well represented in the dendrogram’s right wing, with Flann O’Brien, James Joyce, Samuel Beckett and Djuna Barnes having clustered together, though I am not quite sure what to make of Ford Madox Ford/Joseph Conrad’s showing at all, being unfamiliar with the work.

IV: Conclusion

The basic rule in interpreting dendrograms is that the closer the ‘leaves’ reach the bottom, the more similar they can be said to be. Therefore, Anne Enright and Will Self are the contemporary modernists most closely aligned to the forebears, if indeed forebears they can be said to be. It would be harder, from a quantitative perspective, to align Sara Baume with this trend in a straightforward manner, and McBride only seems to correlate with Stein because of how inalienably strange their respective prose styles are.

The primary point to take away here, if there is one, is that more investigations are required. The analysis is hardly unproblematic. For one, the corpus sizes vary enormously. Borges’ corpus is around 46 thousand words, whereas Proust reaches somewhere around 1.2 million. In one way, the results are encouraging, Borges and Barnes, two authors with only one texts in their corpus, aren’t prevented from being compared to novelists with serious word counts, but in another way, it is pretty well impossible to derive literary measurements from texts without taking their length into account. The next stage of the analysis will probably involve breaking the corpora up into units of 50 thousand words, so that the results for individual novels can be compared.

Re-reading Eimear McBride’s ‘A Girl is a Half-Formed Thing’

A book that I’m looking forward to reading, that doesn’t exist yet, is an academic account of how Irish contemporary fiction went, in such a short space of time, from social realism, to the precociously sentenced art writing with dissociative narrators that now composes the Irish literary milieu. It’s the sort of thing that was probably brewing for a long time, these trends tend to be, but I first became aware of it when Eimear McBride’s A Girl is a Half-Formed Thing was published in 2013. It caused a bit of stir in the literary press at the time, for its supposed uncompromising experimentalism, and its fraught, J.K. Rowling-esque publication history. Critics compared it to Marcel Proust or Samuel Beckett, but I don’t think there was a single review that didn’t mention James Joyce.

In the works of Sara Baume, Joanna Walsh or Claire-Louise Bennett, there are certainly comparisons to be made along these lines, but I think McBride is the novelist of the current generation who is suffering most egregiously under these comparisons. This leads to a kind of distortion that McBride has spoken about recently, saying that it’s ‘a way of not being seen’. Claire Lowdon, writing on McBride’s prose style in Areté, has used the Joyce comparisons as a way of demeaning the novel’s experimental qualities, saying that they are ‘redundant’ and ‘artificial’:

Having invoked Joyce, Joyce has to be McBride’s standard. She has taken all the difficulty and none of the brilliance.

Lowdon’s reading is important and thorough, but I have problems with it. The most significant one being that I think it’s nonsensical to say that just because a work is in some way formally indebted to Joyce has to be 1) as good, 2) as innovative and 3) as good and as innovative in exactly the same ways. I think it’s a very strange point to make that we should benchmark a writer relative to their influences , particularly when this is a comparison furthered more by the laziness of critics than something that McBride has taken upon herself. It’s also inadequate to assume McBride and Joyce’s modernisms are coterminous; I happen to think that they’re rather distinct in a number of significant ways.

Firstly, it’s clear that A Girl is more formally aligned with the Wake than with Ulysses, but taken relative to the former, A Girl manifests far less attention to the materiality of language. In A Girl, there’s less puns, there’s less references, there’s less leitmotifs. It’s also possible to make sense of A Girl without reference to other works. But it’s a mistake to regard this as McBride’s failure to live up to her twentieth century modernist aesthetics. An example from the novel’s opening that Lowdon cites reads as follows:

For you. You’ll soon. You’ll give her name. In the stitches of her skin she’ll wear your say. Mammy me? Yes you. Bounce the bed I’d say. I’d say that’s what you did. Then lay you down. They cut you round. Wait and hour and day.

‘Wait and hour and day’, carries with it the vague association with the phrase ‘a year and a day’ but it doesn’t strictly make sense in that context, there’s no clear reason for the semantic distortion. But there’s also no requirement that there is, nor that it add up to some enormous mythic framework in the same way that the Wake does. I think that once we approach the novel from this position, one which takes account of McBride’s actual concerns, we’ll be able to come to a more sophisticated understanding that doesn’t amount to downgrading her because of her perceived inadequacy in relation to Joyce.

By her own admission McBride retains an interest in nineteenth century novels with less self-consciousness about their language or processes of meaning-making. She has cited the work of the Russian novelist Fyodor Dostoevsky as significant, particularly as an example of proto-modernism, or modernism in a nascent stage of its development, wherein human intersubjectivity was beginning to make itself known within the novel while the tenets of realistic fiction was still trying to accommodate it. Being aware of the fact that The Lesser Bohemians is not the novel under discussion, it’s important to note the way in which it demonstrates this interplay. Within the context of what has been referred to by the author as a ‘modernist monologue’ there is a very sensationalistic narrative in which a character lays out their life story in a very direct and straightforward manner in the same way that you might find extended and directly rendered narratives nested within nineteenth century novels. McBride has said that this is a very deliberate formal mechanic which is pertinent to the text’s thematic concerns, as it is a novel about relating to another person in spite of one’s traumatic past:

In the end you tell a person and you have to use the words that they’ll understand.

What makes McBride’s modernism distinct then, is the centrality it gives to the conveying of narrative information, deploying it as a means of bringing the reader closer to

physical experience, to write about the female experience…the reader can partake in the experience.

McBride has said that the language of A Girl, was written in a way that would create a physical experience for the reader, an immediacy on the page that is reminiscent of theatre. She’s expressed frustration at the content of many of her reviews which have emphasised the quality of the language at the expense of the novel’s content, which she regards as very significant. This stands in contrast to the tradition of the Wake or other modernist works famed for their unintelligibility, such as Gertrude Stein’s The Making of Americans: Being a History of a Family’s Progress is a novel that she has spoken about dismissively for being ‘too navel-gaze-y.’

This stated interest in what the book is ‘about’ and a reader-centric ethic, is I think at least a partial reversal of expectations within the modernist tradition. McBride’s modernism is therefore conceptualised, not as a constructed textual estrangement from reality, but an attempt to bring it closer, to a dwelling-place of authentic being. Not that it’s likely to close off such comparisons in the future.

Can a recurrent neural network write good prose?

At this stage in my PhD research into literary style I am looking to machine learning and neural networks, and moving away from stylostatistical methodologies, partially out of fatigue. Statistical analyses are intensely process-based and always open, it seems to me, to fairly egregious ‘nudging’ in the name of reaching favourable outcomes. This brings a kind of bathos to some statistical analyses, as they account, for a greater extent than I’d like, for methodology and process, with the result that the novelty these approaches might have brought us are neglected. I have nothing against this emphasis on process necessarily, but I do also have a thing for outcomes, as well as the mysticism and relativity machine learning can bring, alienating us as it does from the process of the script’s decision making.

I first heard of the sci-fi writer from a colleague of mine in my department. It’s Robin Sloan’s plug-in for the script-writing interface Atom which allows you to ‘autocomplete’ texts based on your input. After sixteen hours of installing, uninstalling, moving directories around and looking up stackoverflow, I got it to work.I typed in some Joyce and got stuff about Chinese spaceships as output, which was great, but science fiction isn’t exactly my area, and I wanted to train the network on a corpus of modernist fiction. Fortunately, I had the complete works of Joyce, Virginia Woolf, Gertrude Stein, Sara Baume, Anne Enright, Will Self, F. Scott FitzGerald, Eimear McBride, Ernest Hemingway, Jorge Luis Borges, Joseph Conrad, Ford Madox Ford, Franz Kafka, Katherine Mansfield, Marcel Proust, Elizabeth Bowen, Samuel Beckett, Flann O’Brien, Djuna Barnes, William Faulkner & D.H. Lawrence to hand.

My understanding of this recurrent neural network, such as it is, runs as follows. The script reads the entire corpus of over 100 novels, and calculates the distance that separates every word from every other word. The network then hazards a guess as to what word follows the word or words that you present it with, then validates this against what its actuality. It then does so over and over and over, getting ‘better’ at predicting each time. The size of the corpus is significant in determining the length of time this will take, and mine required something around twelve days. I had to cut it off after twenty four hours because I was afraid my laptop wouldn’t be able to handle it. At this point it had carried out the process 135000 times, just below 10% of the full process. Once I get access to a computer with better hardware I can look into getting better results.

How this will feed into my thesis remains nebulous, I might move in a sociological direction and take survey data on how close they reckon the final result approximates literary prose. But at this point I’m interested in what impact it might conceivably have on my own writing. I am currently trying to sustain progress on my first novel alongside my research, so, in a self-interested enough way, I pose the question, can neural networks be used in the creation of good prose?

There have been many books written on the place of cliometric methodologies in literary history. I’m thinking here of William S. Burroughs’ cut-ups, Mallarmé’s infinite book of sonnets, and the brief flirtation the literary world had with hypertext in the 90’s, but beyond of the avant-garde, I don’t think I could think of an example of an author who has foregrounded their use of numerical methods of composition. A poet friend of mine has dabbled in this sort of thing but finds it expedient to not emphasise the aleatory aspect of what she’s doing, as publishers tend to give a frosty reception when their writers suggest that their work is automated to some extent.

And I can see where they’re coming from. No matter how good they get at it, I’m unlikely to get to a point where I’ll read automatically generated literary art. Speaking for myself, when I’m reading, it is not just about the words. I’m reading Enright or Woolf or Pynchon because I’m as interested in them as I am in what they produce. How synthetic would it be to set Faulkner and McCarthy in conversation with one another if their congruencies were wholly manufactured by outside interpretation or an anonymous algorithmic process as opposed to the discursive tissue of literary sphere, if a work didn’t arise from material and actual conditions? I know I’m making a lot of value-based assessments here that wouldn’t have a place in academic discourse, and on that basis what I’m saying is indefensible, but the probabilistic infinitude of it bothers me too. When I think about all the novelists I have yet to read I immediately get panicky about my own death, and the limitless possibilities of neural networks to churn out tomes and tomes of literary data in seconds just seems to me to exacerbate the problem.

However, speaking outside of my reader-identity, as a writer, I find it invigorating. My biggest problem as a writer isn’t writing nice sentences, given enough time I’m more than capable of that, the difficulty is finding things to wrap them around. Mood, tone, image, aren’t daunting, but a text’s momentum, the plot, I suppose, eludes me completely. It’s not something that bothers me, I consider plot to be a necessary evil, and resent novels that suspend information in a deliberate, keep-you-on-the-hook sort of way, but the ‘what next’ of composition is still a knotty issue.

The generation of text could be a useful way of getting an intelligent prompt that stylistically ‘borrows’ from a broad base of literary data, smashing words and images together in a generative manner to get the associative faculties going. I’m not suggesting that these scripts would be successful were they autonomous, I think we’re a few years off one of these algorithms writing a good novel, but I hope to demonstrate that my circa 350 generated words would be successful in facilitating the process of composition:

be as the whoo, put out and going to Ingleway effect themselves old shadows as she was like a farmers of his lake, for all or grips — that else bigs they perfectly clothes and the table and chest and under her destynets called a fingers of hanged staircase and cropping in her hand from him, “never married them my said?” know’s prode another hold of the utals of the bright silence and now he was much renderuched, his eyes. It was her natural dependent clothes, cattle that they came in loads of the remarks he was there inside him. There were she was solid drugs.

“I’m sons to see, then?’ she have no such description. The legs that somewhere to chair followed, the year disappeared curl at an entire of him frwented her in courage had approached. It was a long rose of visit. The moment, the audience on the people still the gulsion rowed because it was a travalious. But nothing in the rash.

“No, Jane. What does then they all get out him, but? Or perfect?”

“The advices?”

Of came the great as prayer. He said the aspect who, she lay on the white big remarking through the father — of the grandfather did he had seen her engoors, came garden, the irony opposition on his colling of the roof. Next parapes he had coming broken as though they fould

has a sort. Quite angry to captraita in the fact terror, and a sound and then raised the powerful knocking door crawling for a greatly keep, and is so many adventored and men. He went on. He had been her she had happened his hands on a little hand of a letter and a road that he had possibly became childish limp, her keep mind over her face went in himself voice. He came to the table, to a rashes right repairing that he fulfe, but it was soldier, to different and stuff was. The knees as it was a reason and that prone, the soul? And with grikening game. In such an inquisilled-road and commanded for a magbecross that has been deskled, tight gratulations in front standing again, very unrediction and automatiled spench and six in command, a

I don’t think I’d be alone in thinking that there’s some merit in parts of this writing. I wonder if there’s an extent to which Finnegans Wake has ‘tainted’ the corpus somewhat, because stylistically, I think that’s the closest analogue to what could be said to be going on here. Interestingly, it seems to be formulating its own puns, words like ‘unrediction,’ ‘automatiled spench’ (a tantalising meta-textual reference I think) and ‘destynets’, I think, would all be reminiscent of what you could expect to find in any given section of the Wake, but they don’t turn up in the corpus proper, at least according to a ctrl + f search. What this suggests to me is that the algorithm is plotting relationships on the level of the character, as well as phrasal units. However, I don’t recall the sci-fi model turning up paragraphs that were quite so disjointed and surreal — they didn’t make loads of sense, but they were recognisable, as grammatically coherent chunks of text. Although this could be the result of working with a partially trained model.

So, how might they feed our creative process? Here’s my attempt at making nice sentences out of the above.

— I have never been married, she said. — There’s no good to be gotten out of that sort of thing at all.

He’d use his hands to do chin-ups, pull himself up over the second staircase that hung over the landing, and he’d hang then, wriggling across the awning it created over the first set of stairs, grunting out eight to ten numbers each time he passed, his feet just missing the carpeted surface of the real stairs, the proper stairs.

Every time she walked between them she would wonder which of the two that she preferred. Not the one that she preferred, but the one that were more her, which one of these two am I, which one of these two is actually me? It was the feeling of moving between the two that she could remember, not his hands. They were just an afterthought, something cropped in in retrospect.

She can’t remember her sons either.

Her life had been a slow rise, to come to what it was. A house full of men, chairs and staircases, and she wished for it now to coil into itself, like the corners of stale newspapers.

The first thing you’ll notice about this is that it is a lot shorter. I started off by traducing the above, in as much as possible, into ‘plain words’ while remaining faithful to the n-grams I liked, like ‘bright silence’ ‘old shadows’ and ‘great as prayer’. In order to create images that play off one another, and to account for the dialogue, sentences that seemed to be doing similar things began to cluster together, so paragraphs organically started to shrink. Ultimately, once the ‘purpose’ of what I was doing started to come out, a critique of bourgeois values, memory loss, the nice phrasal units started to become spurious, and the eight or so paragraphs collapsed into the three and a half above. This is also ones of my biggest writing issues, I’ll type three full pages and after the editing process they’ll come to no more than 1.5 paragraphs, maybe?

The thematic sense of dislocation and fragmentation could be a product of the source material, but most things I write are about substance-abusing depressives with broken brains cos I’m a twenty-five year old petit-bourgeois male. There’s also a fairly pallid Enright vibe to what I’ve done with the above, I think the staircases line could come straight out of The Portable Virgin.

Maybe a more well-trained corpus could provide better prompts, but overall, if you want better results out of this for any kind of creative praxis, it’s probably better to be a good writer.