Wednesday, February 25, 2009

reduplication

I don't know how well your average US movie trailer represents spoken North American English, but I found it rather striking that this particular trailer contained not one, but two examples of contrastive focus reduplication in the space of 40 seconds. The first one is a classic:

00:43-00:46
- You like her?
- Yes.
- You LIKE-HER-like-her?

It's the second one, spoken by the lovely and talented Zooey Deschanel, that stands out:

01:20-01:22
- I got the baby.
- THE-BABY-the-baby?

According to Ghomeshi et al. referenced in the LanguageLog post above (preprint, section 3.4, p. 27), one of the constraints on contrastive reduplication (CR) is

(57) a. The scope of CR is either X0 or XPmin.

So if I got this right, in case of noun phrases, only minimal noun phrases (i.e. bare noun/pronoun) can be reduplicated and thus reduplication of a NP -> D N such as the one above should be impossible. And indeed, as Ghomeshi et al. elaborate:

Condition (57a) is violated in *A-LINGUIST-a linguist (cf. (47)): although the determiner is a grammatical morpheme, it is outside of NPmin, so it cannot be within the scope of CR.

The examples they refer to are

(47) a. Do you want [tu:] or WANT [tu:]-want [tu:]?
b. ? Do you wanna or WANNA-wanna?

(48) a. I wouldn’t DATE–date a linguist.
b. * I wouldn’t DATE-A–date-a linguist

and

(49) * I wouldn't date [CG A-LINGUIST]-[a linguist]

all from section 3.3 which deals with the optionality of some types of complements with some types of heads (e.g. pronouns or prepositions for verbs, PPs for adjectives). In comments to (47), they wonder if this optionality can be explained phonologically, i.e. those complements are considered clitics and they can, but don't have to be, included in CR. Long story short, citing Hayes' "clitic group" (CG in (49) - "a prosodic word plus the clitics to its right or left"), Ghomeshi et al. dismiss the idea of clitics being involved in CR, noting that (emphasis mine)

So in (48), date and a do not form a copyable clitic group, since a must belong to the same clitic group as linguist. However, Hayes’ definition of clitic group includes clitics on the left as well as the right, and these never reduplicate11.

Hence the asterisk preceding (49) - whatever the phonological considerations in CR, clitics to the left of the prosodic word (such as determiners in NPs) are never a part of the duplicated phrase. Footnote 11 explains that the only exception to that rule are lexicalized proper names (The Hague) and gives the following example:

The casino isn't in THE-PAS-The-Pas, but in Opaskwayak.
(The reserve of the Opaskwayak Cree Nation includes land that is in the northern Manitoba town of The Pas, but not legally part of it.)

This example, however, does not appear in Kevin Russell's corpus and, more importantly, it doesn't explain Ms. Deschanel's reduplication.

Now I'm way in over my head here, so my first stupid question is whether the fact that "THE-BABY-the-baby" doesn't appear in a complete sentence is to any extent relevant. Judging by the examples of same type given by Ghomeshi et al. (e.g. 7, 8, 10, 11, 17), I don't think so, but what do I know. The bottom line seems to be that we have a real-life example of reduplication of a NP -> D N, something Ghomeshi et al. claim is not possible. Would you agree or did I miss something?

Monday, January 12, 2009

€

Now that the first working week of the new year is over, I think it is safe to say that Slovakia's transition to the new currency is going very well. I was a little late on board, making my first official euro purchase on Thursday (shoelaces, € 0.83) and making the first ATM withdrawal on Friday (€ 40), but even though I am arithmetically challenged, so far so good.

Unlike other nations - such as Malta - we were fortunate enough to avoid major linguistic issues associated with the adoption of the euro, but there are still some minor changes to consider. I, for one, rejoice at the thought of never having to decide again whether I should translate "slovenská koruna" as "Slovak crown" (which seems to have been the preferred form) or "Slovak koruna" (which sounds better to me). But that's just a minor point. As some, including this report in Pravda and this one in SME (originally by the Czech press agency ČTK), have pointed out, the real story is the changes which euro will mean for the slang terms for amounts, coins and banknotes:

desiatka/desík/desina = 10 (note the typical Bratislava suffix -ina normally used in standard Slovak to form fractions)
dvacka = 20 (but not dvacina)
pajdík = 50 (another typical Bratislava / Western Slovak formation)
kilo = 100
pětikilo = 500
liter = 1000
melón = 1 000 000 (melón also meaning, of course, "melon", especially "water melon")

But just what kind of changes should we expect? According to the Pravda editorial,

Slangové označenia peňazí ako desík, kilo, liter či melón odídu s korunou do zabudnutia. ... Liter stratí s eurom zmysel úplne. Tisíceurová bankovka totiž neexistuje, najvyššie papierové euro má hodnotu 500.

Slang terms designating money like desík, kilo, liter or melón will become obsolete. With the adoption of the euro, liter will vanish completely. There is no € 1000 note, the highest demonination is € 500.

Really? Does that mean that we will no longer count money in tens, hundreds, thousands or millions? I guess it would be absolutely futile to try to explain to a journalist that one single word can be used for an object, like a banknote, as well as a concept, like an amount. Martin Považaj, a linguist with the Slovak Academy of Sciences, did try to do so when talking to ČTK with questionable results:

"Je však možné, že niektoré z týchto slov zostanú, ale nadobudnú novú významovú náplň, to znamená, že číselná hodnota skrývajúca sa za týmito slovami zostane, ale bude sa už vzťahovať na eurá[.]"

"It is possible that some of these words will stay with us, but will acquire a new meaning, that is the numerical value behind the words will remain, but will refer to amounts in euro."

So according to Dr. Považaj, it is merely possible that with a change in extralinguistic reality, the language will follow suit. Whereas according to anybody else with a bit of understanding of language, it is, how should I put it, pretty fucking certain.
And then the author of the report helpfully adds:

Znamenalo by to, že tieto výrazy by vyjadrovali 30-krát väčšiu hodnotu, kilo by tak už napríklad nebolo 100 Sk (3,32 eura), ale sto eur (3013 Sk).

This would mean that these terms would be used for amounts 30 times higher that before, thus kilo would not mean 100 Sk (3.32 euro), but hundred euro (3013 Sk).

The fact that this needs to be spelled out astonishes me. I'm quite certain that this - together with Dr. Považaj's explanation above - is a statement to the view of Slovak as something rigid and immutable so prevalent in our society. If you want another example, just try the very next paragraph:

Slová ako päťeurovka alebo stoeurovka sa doteraz do slovníkov slovenského jazyka nedostali, podľa jazykovedcov sú však spisovné a využívajú sa v hovorovej neoficiálnej komunikácii.

Words like päťeurovka (5 € note) or stoeurovka (100 € note) have not yet been included in dictionaries of Slovak, but according to the linguists, these are standard terms which are used in spoken unofficial communication.

Two perfectly legit compounds made from two perfectly normal (and standard) words based on a long-used and perfectly standard terms (päťkorunáčka and stokorunáčka) and yet people still feel the need to ask the official body to please please pretty please validate their own words. This makes me glad we haven't had to deal with any serious linguistic issues. Although the following clusterfuck could have been pretty funny to watch...

And finally, there is the issue of the new slang term for euro. According to Mira Nábělková, a Czech linguist quoted by the SME/ČTK piece, terms like jurko, jurášky, juráše and juroše have been recorded on the internet, obviously combining the English pronunciation with Slovak suffixes. I can see why jurko (note the diminutive suffix -ko) would work, but since it happens to be the diminutive form of the name Juraj, I don't think it's very likely. As for jurášky, I only found one single occurrence, and that one insists it's a pronunciation used by speakers of English. There were a few ghits on juráše, but all of those were from websites in Czech and as for juroše, that one only appears in variations of this ČTK report. And to add one final insult, Dr. Nábělková immediately connects these imaginary slang terms for euro with Juraj Jánošík, i.e. the most stereotypical stereotype in the history of Slovak stereotypes. Even her other examples, the diminutives eurko (neuter), eurík (masculine) and eurka (feminine) seem fishy. A brief Google search quickly revealed that the feminine form is nothing of the kind, but rather Nominative plural of the neuter form (UPDATE: but only when written without diacritics, the proper plural form is eurká). Examples (the first three ghits):

1. ... na ktore mimochodom sa eurka uz vyfasovali ... (which, by the way, they already got euros for)
2. ... a nosi eurka za kazdy mliecny zubok ... (who brings euros for every milk tooth)
3. ... a ked majitelom Slovanu dojdu eurka ... (and when the proprietors of Slovan run out of euros)

From the final list of terms collected on the internet, eurčeky is another nonce formation, but euráče, euráky, euráčiky (another diminutive) and - much less common - euroše are in actual use with euráče being the most common, at least according to raw ghits (2460 vs. 116, 224 and 109). I guess only time will tell which one(s) will be left standing. But if you want to come back in a few months and find out, be sure not to rely on Pravda, SME or ČTK.

Tuesday, October 21, 2008

door

In what is just another episode of a long-running series, today I was once again forced to deal with medical professionals. That's usually bad enough - doctors routinely make the top of my shit list with nurses right behind. What made it worse is that instead of going to my usual place, a rather friendly clinic in a convenient location situated next to a lovely park especially beautiful this time of year, I had to drag myself over to this butt-ugly God-forsaken communist-era hospital complex on the outskirts of town. Long story short, I wasted about three hours, didn't even get to see the doctor and most likely caught something along the way. Not a good day, if you catch my drift. All would have been lost, had I not stumbled across this while I wandered the halls:


Bilingual (Slovak-English) and trilingual (Slovak-English-German) signs are not unusual in Bratislava - in fact, the aforementioned rather friendly clinic employs them routinely, considering the large number of foreign residents and Austrians who either live here or come here to get high-quality medical care (especially dental) at a very low cost. But this is the first time I've seen the ubiquitous (in hospitals and clinics, that is) "Don't knock" sign translated into Chinese. Why did this particular immunology clinic put up this sign, I don't know. Bratislava does have a relatively large Chinese community, yet somehow I doubt its members are particularly susceptible to alergies and autoimmune diseases, considering that this was the only door with a sign in Chinese.

Be that as it may, I naturally had to check if the Chinese was legit. Of the five characters 请不要敲门, I only recognized the negative particle 不 bù. The rest was supplied by various dictionaries:

  • 请 [qǐng] = to ask, to invite, please
  • 不 [bù] = (negative prefix), not, no
  • 要 [yào] = important, to want, FUT AUX, may, must, OR
  • 要 [yāo] = demand, ask, request, coerce
  • 敲 [qiāo] = knock, to strike, to knock (at a door), to hit
  • 门 [mén] = gate, door

Seeing as I get about 4000 hits when googling the phrase in quotes and Google Translate provides this very phrase as the translation of "Please don't knock", I suspect it's indeed proper Chinese. Please don't hesitate to correct me if I'm wrong. That would certainly be interesting, just consider Engrish. With China playing an ever increasing economic and political role on the global stage, Chinese is bound to increase in importance and stature and will inevitably be used by people as clueless about it as the authors of the many Engrish texts are about English. Is it possible that I have just witnessed the birth of Hanyish or perhaps Zhongwenish?

Sunday, September 28, 2008

google

As you probably already know, Google Translate has added 11 more languages, including Slovak, to its already impressive portfolio. While testing the English-Slovak service, I was pleasantly surprised at the MT engine's ability to handle syntax like noun phrases containing adjectives, although I noted a number of problems associated with translating English idiomatic structures, such as those involving verbs "give" and "take" or multiword expressions. Overall, about 60% of translations of work related documents I put in did produce comprehensible and usable texts, so color me impressed. More testing will be required to see if Slovak translators who work for me should start to worry about their jobs (and trust me, I do have a shit list), but I'm pleased to inform you that we already have a candidate for the mistranslation of the year. Consider the headline of this report on the first US presidential debate from CNN.com and then have a look at the translation, especially the items in red:



English: Analysis: A few jabs, but no knockout in first debate
Slovak: Analýza: Za pár popíchnutí, ale žiadna kokot nedved v prvom diskusie

OK, my praise of syntax handling now sounds premature, since in "v prvom diskusie", neither the noun nor the ordinal numeral are declined properly (it should be "v prvej diskusii"), but that's not the interesting bit. That rests with the translation of the word "knockout": "kokot nedved". "Kokot" = "dick, prick" is of course the basic Slovak insult for a man, for more information see here. It is also a very vulgar term, rarely seen in print or heard on the airwaves, so its appearance here will not only ellicit a chuckle for its own sake, but also the question of just what corpus was Google using in training the MT engine. The web, sure, but I can't think of any sufficiently large bilingual corpus where that word would crop up. And that question is even more justified with the second part of the translation: what the hell is a "nedved"? The only word that even comes close is the Czech surname Nedvěd which is a form of "medvěd" = "bear". There are a few people with that name with a significance presence on the web to be included in a web corpus, like the football player Pavel Nedvěd, the hockey player Petr Nedvěd and the folk singers Jan (Honza) and František Nedvěd. But how did their name get into the translation for "knockout"? "Knockout" ("knokaut" in Slovak) is a sports term, but I know of no boxer by the name of Nedvěd. Then again, football and hockey players as fellow athletes could probably fit the bill. That still leaves the question of how did this Czech word get into a translation into Slovak. And it's not the only one - if you look at the screenshot, you will see at least three more words clearly identifiable as Czech (highlighted in green):

- "štípnout" for "tweak" - Slovak: "upraviť, vyladiť"; "štípnout" = "pinch, sting", slang: "steal", although one of my dictionaries gives "štípnout" for tweak" without any further context or explanation.
- "slíbený" for "vowed" - Slovak: "sľúbený". Note that this is a past participle while the original has "vowed" as a past tense verb.
- "poldové" for "cops" - Slovak: "policajti", slang: "fízli". Note the context mismatch: both "poldové" and "fízli" is stylistically marked and not very likely to appear in a newspaper save perhaps for direct quotes.

Once again, I assume that web corpora were to some extent used to train the MT engine. As my own feeble attempts at corpus research have shown, the country code cz or sk in the domain name does in no way guarantee that you will find only Czech or Slovak text there. The actual ratio is hard to determine, but it is definitely nice to see that one of the better aspects of Czechoslovakia - its almost fully bilingual citizens - survives to this day.

And one last interesting bit from this small test: Barack Obama's full name is translated as it should be. But whenever his last name shows up on its own, Google translates it as "osobách" = "person-LOC.PL" (highlighted in light blue). Buggered if I know why...

(h/t: filer)

Monday, August 04, 2008

awwissu

In an effort to promote and protect the Maltese language and establish an unified linguistic policy, in 2005, the government and parliament of Malta have adopted the Att dwar l-Ilsien Malti (Maltese Language Act, Chapter 470). The Act establishes Il-Kunsill Nazzjonali tal-Malti (National Council for the Maltese Language) as the main body charged with "adopting and promoting a suitable language policy and strategy" (Part II, 4(1)). Aside from the general task of promoting Maltese Language both in Malta and abroad (Part II, 5(1)), the Council has also been specifically charged with updating "the orthography of the Maltese Language as necessary" and establishing "the correct manner of writing words and phrases which enter the Maltese Language from other tongues" (Part II, 5(2)). Having been formed in 2005, the Council, headed by prof. Manwel Mifsud, immediately began working on an orthography reform and three years of research, public debate and expert discussion resulted in the publication of Government notice no. 642 in the Government Gazette of July 25th which amends the official orthography of Maltese. This amendment, also known as Deċiżjonijiet 1, is the third official update of Maltese orthography established by the Tagħrif fuq il-Kitba Maltija written by Ninu Cremona and Ġanni Vassallo and published in 1924 by Akkademja tal-Malti. Unlike the previous ones, Żieda mat-Tagħrif of 1984 and Aġġornament tat-Tagħrif fuq il-Kitba Maltija of 1992, this reform will consist of three parts.

As outlined by the guide to the decision-making process published by the Council under the title "It-Triq lejn id-Deċiżjonijiet 1" (8 MB PowerPoint presentation), the Council identified three problematic areas:

1. Orthographic variants
2. English loans
3. Phonetic variants

A decision was made to deal with these issues one by one in this particular order. In the first phase, over 300 orthographic variants were collected, after which the Council issued a general call for opinions to the public and a specific one to a selected group of professionals (authors, translators, teachers and journalists). It is interesting to note that the latter call was answered by only 35 people, fortunately including some of the biggest names in Maltese literature. The reactions were published in a separate volume of 195 pages titled Innaqqsu l-inċertezzi. The public debate was concluded by a workshop on orthographic variants with over 180 participants and Innaqqsu l-inċertezzi as the main subject of discussion. The final decision was entrusted to a committee chaired by Albert Borg consisting of 11 experts. After 30 meetings, the committee issued a final recommendation which was unanimously approved by the Council and finally published as Government notice no. 642, a document with the legal force of a law entering into effect on July 25th, 2008.

Government notice no. 642 / Deċiżjonijiet 1 consists of five main sections:

1. Grave accent and circumflex
Only grave accent is now used in Maltese (e.g. kafè, però), circumflex is abolished.

2. Capitalization
Deċiżjonijiet 1 establishes comprehensive rules for capitalization and lack thereof. Most notably, names of religions, religious orders, sects, art movements, styles and epochs as well as adjectives derived from them and names of their members are now capitalized throughout, e.g. l-Iżlam, ir-Rinaxximent, l-Impresijonisti, stil Sikulo-Normann, id-Dumnikani and in-Nazzjonalisti.

3. Word combinations
This section deals with various phrases, fixed expressions and idioms where there's been significant confusion. Subsection 3.1 covers expresions and idioms where the constituent parts are written separately, such as the reduplicative constructions of the type ftit ftit (little by little, gradually), numerals except 11-19, adverb nett and preposition a la and għala.
Subsection 3.2 contains rules for writing phrases and expressions written together, separately or hyphenated. 3.2.1 covers (mostly prepositional) phrases which can be written together or joined by a hyphen, such as fil-waqt (at the time of) as opposed to filwaqt (while). Appendix A provides a comprehensive list of such multiword expressions that are now written as one word, appendix B contains those that are still written separately or hyphenated.
3.2.2. deals with prefixes such as awto-, ko-, anti- and post- which are now written without a hyphen except where the stem is capitalized, e.g. antiinflammatorju (or antiflammatorju, both are acceptable), but anti-Iżlamiku.
The rest of subsection 3.2 covers prepositions ġo, ma', sa' and ta' (actually the Genitive exponent), the negative particle ma and the preposition kontra. New rules provide a choice between writing ma, ma', sa' and ta' in both full and short form (sa issa / s'issa "until now") when followed by a vowel, għ or ħ. Subsection 3.2 also significantly simplifies the spelling of ġo, ma', sa' and ta' + definite article il-. Forms ġol, mal, sal and tal are now used regardless of what follows (consonant, vowel, għ or ħ), removing one major headache for all speakers of Maltese.

4. Roots and stems
This short section establishes different rules for writing words of Semitic origin and Romance words. For Semitic roots, the rules confirm the practice of using the same set of letters for one root even though the pronunciation may differ (e.g. ktibna "we wrote" is pronounced [ktibna], but kitbu "they wrote" is pronounced [kidbu]). The obvious exceptions with some roots, such as verbae tertiae għ, still apply (e.g. the verb sema' with the root smgħ and first person singular perfect smajt). Non-Semitic stems are written the way they occur in words adopted into Maltese and no effort is made to establish a single correct spelling.

5. Other
This final section attempts to simplify and unify the orthography of a small number of words. Subsection 5.1 deals with consonants with the same pronunciation, but non-phonetic spelling or two different spellings, where in both cases the phonetic spelling is chosen as the preferred variant. Examples include dvalja which replaces tvalja "tablecloth", risq instead of riżq "profit, benefit" and skont instead of skond "according to" which was even before the reform written with t before enclictic pronouns. Subsection 5.2 also adapts the spelling of some words to match their most common current pronunciation, such as the title of this post Awwissu instead of Awissu "August", dettall instead of dettal, Iżlam instead of Islam and prefers karozza over of karrozza "car". The rest of the section covers a number of words designed to bring their spelling in line with their etymology (Magreb instead of Maghreb, since the word was borrowed from Italian or English) and three spelling changes based on reinterpretation of roots.

Having been published in the Government Gazette, the new rules of Maltese Orthography are now binding for all government institutions including schools, textbooks and examinations. The Government notice no. 642, however, provides for a three-year transitional period during which both variants will be acceptable. On July 25th, 2011, the new forms will finally become the only correct and acceptable ones. It will remain to be seen how speakers of Maltese will get used to it. We'll see if the first reactions were an indicator and if, what will that mean for the second phase which deals with borrowings from English and which is well underway. Should be interesting to watch.

UPDATE: Albert Borg comments on the process and the motivations in an interview for Times of Malta published today.