Thursday, June 11, 2009

Special Old English characters on the Mac

Make sure you are using the US Extended keyboard layout:

æALT+'
ÆALT+SHIFT+'
þALT+t
ÞALT+SHIFT+t
ðALT+d
ÐALT+SHIFT+d
āALT+a a
ēALT+a e
...ALT+a ...
ċALT+w c
ġALT+w g
ĊALT+w C
ĠALT+w G

Thursday swimming

60 lengths (1500 yards) in 40 minutes.

Wednesday, June 10, 2009

Old English noun declension: strong nouns

Masculine (stone):

singularplural
nomstānstānas
accstānstānas
genstānesstāna
datstānestānum

Neuter short (ship):

singularplural
nomscipscipu
accscipscipu
genscipesscipa
datscipescipum

Neuter long (word):

singularplural
nomwordword
accwordword
genwordesworda
datwordewordum

Feminine short (gift):

singularplural
nomġiefu [v]ġiefa
accġiefeġiefa
genġiefeġiefa
datġiefeġiefum

Tuesday, June 9, 2009

Midweek swimming

60 lengths (1500 yards) in 40 minutes.

Saturday, June 6, 2009

Swimming

60 lengths (1500 yards) in 40 minutes.

Thursday, June 4, 2009

OpenCCG 'morph.xml' files

A morph.xml file is a list of word-forms and a list of associated macros.

morph @name
  > entry* @word @stem @pos @macros
  > macro* @name
      > fs? @id @attr @val
      > lf?

The 'stem' attribute should be thought of as the semantic predicate (by default it's the same as the value of 'word'). The 'pos' attribute specifies which families are relevant in building the full lexical category. The 'id' attribute make reference to an identifier in the 'lexicon.xml' file.

Simple Old English words (monosyllables)

bedbɛdbed
blódblo:dblood
GodgɔdGod
landlandland
nimnɪmtake
rammramram
beby
nannannone/no
tídti:dtime
inɪnin
onɔnon
upʊpup
binbɪnbin
swiftswɪftswift
goldgoldgold
grimgrɪmgrim
songsɔŋgsong
cornkɔrncorn
lamblamblamb
windwɪndwind
þusθʊsthus
farfarfar
eftɛftagain/back
selfsɛlfself
fyrfyrfire
sweordswɛɔrdsword
cræftkræftcraft
léoflɛɔfdear (sir)
fréofrɛɔfree
he:he
himhɪmhim
nihtnɪçtnight
handhandhand
þurhθʊrxthrough
icɪtʃI
cwæþkwæ&thetasaid
cildtʃɪldchild
gadgadgoad
græggræjgray
cnihtknɪçtboy/servant
betstbɛtstbest
fiscfɪʃfish
scipʃɪpship
ecgɛdʒedge
bricgbrɪdʒbridge
secgsɛdʒman

Wednesday, June 3, 2009

Old English consonants

The following consonants are straightforward:

p [p]t [t]k [k]
b [b]d [d]
m [m]n [n]
l [l]r [r]
x [ks]

Fricatives each have two allophones:

f [f/v]
s [s/z]
þ [θ/ð]
ð [θ/ð]

The voiceless allophone appears initially, finally, and adjacent to a voiceless consonant. The voiced allophone appears between two vowels, or between a vowel and a voiced consonant.

The 'h' sounds has three allophones:

h [h/ç/x]

The [h] sound appears initially, [ç] after a front vowel, and [x] after a back vowel.

The sounds 'c' and 'g' have two allophones:

c [tʃ/k]g [j/g]

The first allophone appears: (a) before a front vowel; and (b) at the end of a word, after a front vowel. The second allophone appears before a back vowel or a consonant.

Some special combinations:

cn [kn]
gn [gn]
sc [ʃ]
cg [dʒ]

Old English vowels

Old English texts contain seven letters which denote vowels. There is evidence that each of these represents two distinct phonemes, one 'short' and one 'long':

majusculeminusculeshortlong
aAa or ɑɑ:
eEɛ or ee:
iIɪ or ii:
oOɔ or oo:
uUʊ or uu:
yYyy:
æÆææ:

Note that there are two different proposed pronunciations for some of the short vowels.

The evidence for phonemic status of vowel length involves the existence of minimal pairs involving homographs like:

writtenshort reflexlong reflex
hamhamhome
isisice
rodrodrood (i.e. cross)

Because one reading of each homograph underwent a vowel change during the Great Vowel Shift, it is believed that these words were not homophones in Old English. In addition, we can use evidence of vowel shift in Modern English to identify long vowels in Old English.

In modern versions of Old English texts, long vowels are often marked with a macron: ā, ē, ī etc. The original Old English texts often use different glyphs: (a) 'i' without a dot; (b) 'y' with a dot; (c) a version of 'E' that looks a bit like 'C'; (d) sometimes a version of 'a' without the curly bit on top.

There are three digraphs/diphthongs, each of which can be either short or long:

digraphsound
eoeo or eʊ
eaæɑ
ieɪ

Note that the digraph 'ie' was probably not pronounced as a diphthong. Long diphthongs/digraphs are conventionally designated by placing a macron above the first letter in the digraph.

Note also that the short vowels 'a', 'æ' and 'ea' are closely related - they evolved out of a single Proto-Germanic vowel 'a' (by phonological conditioning), and evolved back into a single Middle English vowel 'a'. This is important to know to understand certain inflectional paradigms, e.g. 'dæg' vs 'dagas'; 'geat' vs 'gatu'.

Monosyllabic English words with 'ee': A-F

eel
bee
fee
flee
free
knee
beech
beef
been
beep
beer
beet
bleed
bleep
breech
breed
breeze
cheek
cheep
cheer
cheese
deed
deem
deep
deer
feed
feel
freeze

Monosyllabic English words with 'ea': A-M

A * is used when the pronunciations is not /i:/. Generally pronounced /ɛ:/ before the Great Vowel Shift, which rose to /e:/ then /i:/.

each
ear
earl*
earn*
earth*
ease
east
eat
eave

flea

beach
bead
beak
beam
bean
bear*
beard
beast
beat
bleach
bleak
bleat
breach
bread*
breadth*
break*
bream
breast*
breath*
breathe
cease
dead*
deaf*
deal
dean
dear
dearth*
death*
dread*
dream

fear
feast
feat
freak
gear
gleam
glean
grease
great*
head*
heal
health*
heap
hear
heard*
hearse*
heart*
hearth*
heat
heath
heave
jeans
leach
lead[*]
leaf
league
leak
lean
leap
learn*
lease
least
leave
leaves
mead
meal
mean
means
meant*
meat

Others:

feather*
heather*
heaven*
jealous*
leather*
leaven*
meadow*
measure*

The Great Vowel Shift

The long/stressed vowels of Old/Middle English:

front back
i: u:
e: o:
ɛ: ɔ:
 a: 

Typical orthography for these was as follows:

i:y,imine,sight
u:ou,uhouse
e:ee,e,e_esheep, me, mete
o:oo,oboot
ɛ:ea,ebreak, beak
ɔ:oaloaf
a:aa,aname

The Great Vowel Shift involved everything moving up a slot:

  • i: -> ai
  • u: -> au (unless followed by a labial consonant p/b/m)
  • e: -> i:
  • o: -> u:
  • ɛ: -> e: -> i:
  • ɔ: -> o:
  • a: -> ɛ: -> e:

Monday, June 1, 2009

opennlp.ccg.lexicon.MacroItem

macro @name
  > fs*
  > lf
private String name;
private FeatureStructure[] featStrucs;
private LF[] preds;

opennlp.ccg.lexicon.MorphItem

This class parses the entry elements in morph.xml files:

entry @word @pos (@stem) (@class) (@coart) (@macros) (@excluded)

There are the following fields:

private Word surfaceWord;
private Word word;
private Word coartIndexingWord;
private String[] macros;
private String[] excluded;
private boolean coart = false;

The constructor is quite complicated. The value of 'macros' and 'excluded' is straightforward - you just split the value up into string tokens (macro names and excluded strings). The value of 'surfaceWord' comes from passing the value of the word attribute through the tokeniser in some strange way:

surfaceWord = Grammar.theGrammar.lexicon.tokenizer.parseToken(e.getAttributeValue("word"),coart);

The value of 'word' is derived as follows:

word = word.createFullWord(surfaceWord, 
                           e.getAttributeValue("stem"), 
                           e.getAttributeValue("pos"), 
                           null, 
                           e.getAttributeValue("class"));

opennlp.ccg.lexicon.DataItem

This class parses member elements inside family elements inside lexicon.xml files:

member @stem (@pred)

The fields and constructor are obvious:

private String stem;
private String pred;
public DataItem(org.jdom.Element e) { ... }

If there is no 'pred' attribute, its value is the same as 'stem'.

opennlp.ccg.lexicon.EntriesItem

This class is used to parse entry elements from inside family elements in lexicon.xml files:

entry @name (@stem) (@active) (@indexRel)
  > category

The fields for this class are:

private String name;
private String stem;
private Boolean active;
private String indexRel;
private Category cat;
private Family family;

The constructor method is obvious too:

public EntriesItem(org.jdom.Element e, Family f) {
    this.family = f;
    ...
    cat = CatReader.getCat((Element) e.getChildren().get(0));
}

If there is no 'stem' attribute, this field gets the value Lexicon.DEFAULT_VAL, i.e. "[*DEFAULT*]". The default value of 'active' is Boolean.TRUE. If there is no 'indexRel' attribute, then the value of this field comes from that of the family. Obviously, the last line is the most interesting - it takes the first child only and converts it into a category? What if there are other children?

There are the usual 'get' methods, as well as a couple that get attributes of the immediate superfamily, as well a toString() method.

opennlp.ccg.lexicon.Family

This class is used to parse XML family elements in OpenCCG lexicon.xml files:

family @name @pos (@closed) (@indexRel) (@coartRel)
  > entry*
  > member*

The fields of a Family object are thus as follows:

private String name;
private String pos;
private Boolean closed; 
private String indexRel;
private String coartRel;
private EntriesItem[] entries;
private DataItem[] data; //members
private String supertag;

The supertag is formed by removing slash modalities and other minor features from a category (i.e. from the family name). There are get and set accessor methods for each field. The constructor method is exactly as you'd think:

public Family(org.jdom.Element e) { ... }

See also: opennlp.ccg.lexicon.EntriesItem and opennlp.ccg.lexicon.DataItem.

opennlp.ccg.lexicon.Lexicon

The constructor method of the opennlp.ccg.grammar.Grammar class contains the following code:

public final Lexicon lexicon;
lexicon = new Lexicon(this);
lexicon.init(lexiconUrl,morphUrl);

Here we turn to the code for the opennlp.ccg.lexicon.Lexicon class. Here are two basic fields and the constructor:

public final Grammar grammar;
public final Tokenizer tokenizer;

public Lexicon(Grammar g) {
    this.grammar = g;
    this.tokenizer = new DefaultTokenizer();
}

The init method starts as follows, where there are two input parameters, lexiconUrl and morphUrl:

List<Family> lexicon = null;
List<MorphItem> morph = null;
List<MacroItem> macroModel = null;
lexicon = getLexicon(lexiconUrl);
Pair<List<MorphItem>,List<MacroItem>> morphInfo = getMorph(morphUrl);
morph = morphInfo.a;
macroModel = morphInfo.b;

See also opennlp.ccg.lexicon.Family, opennlp.ccg.lexicon.MorphItem, and opennlp.ccg.lexicon.MacroItem.

opennlp.ccg.grammar.Types

This class implements the notion of a (multiple inheritance) hierachy of syntactic types. It is called by the constructor method of the opennlp.ccg.grammar.Grammar class, as documented here. Each Grammar object has an instance field types of class Types. The constructor method for Grammar calls one of two constructor methods to initialise this field: (a) new Types(URL u,Grammar g); or (b) new Types(Grammar g).

The code for the Types makes crucial reference to the opennlp.ccg.unify.SimpleType class (which implements the opennlp.ccg.unify.Unifiable interface).

Apart from that, the code is a bit of a mess, so I'm going to ignore these classes for the moment.

opennlp.ccg.grammar.Grammar

This class encodes the notion of a CCG grammar, essentially a lexicon, a set of rules and a type hierarchy. Its constructor is called by the opennlp.ccg.TextCCG class, as discussed here. Here are the essentials (ignoring XSLT files for transforming from and to LFs, pitch accents, special tokenisers, supertags etc.):

public final Types types;
public final Lexicon lexicon;
public final RuleGroup rules;
private String grammarName = null;
public static Grammar theGrammar; //nasty hack!
Here are the basics of the constructor, which takes a single input parameter url, of class java.net.URL:
theGrammar = this;
SAXBuilder builder = new SAXBuilder();
Document doc = builder.build(url);
Element root = doc.getRootElement();
grammarName = root.getAttributeValue("name");
...
Element typesElt = root.getChild("types");
URL typesUrl;
if (typesElt!=null)
    typesUrl = new URL(url,typesElt.getAttributeValue("file"));
else typesUrl = null;
Element lexiconElt = root.getChild("lexicon");
URL lexiconUrl = new URL(url,lexiconElt.getAttributeValue("file"));
Element morphElt = root.getChild("morphology");
URL morphUrl = new URL(url,morphElt.getAttributeValue("file"));
Element rulesElt = root.getChild("rules");
URL rulesUrl = new URL(url,rulesElt.getAttributeValue("file"));
...
if (typesUrl!=null) types = new Types(typesUrl,this);
else types = new Types(this);
lexicon = new Lexicon(this);
lexicon.init(lexiconUrl,morphUrl); *****
rules = new RuleGroup(rulesUrl,this);
...