Make sure you are using the US Extended keyboard layout:
| æ | ALT+' |
| Æ | ALT+SHIFT+' |
| þ | ALT+t |
| Þ | ALT+SHIFT+t |
| ð | ALT+d |
| Ð | ALT+SHIFT+d |
| ā | ALT+a a |
| ē | ALT+a e |
| ... | ALT+a ... |
| ċ | ALT+w c |
| ġ | ALT+w g |
| Ċ | ALT+w C |
| Ġ | ALT+w G |
Make sure you are using the US Extended keyboard layout:
| æ | ALT+' |
| Æ | ALT+SHIFT+' |
| þ | ALT+t |
| Þ | ALT+SHIFT+t |
| ð | ALT+d |
| Ð | ALT+SHIFT+d |
| ā | ALT+a a |
| ē | ALT+a e |
| ... | ALT+a ... |
| ċ | ALT+w c |
| ġ | ALT+w g |
| Ċ | ALT+w C |
| Ġ | ALT+w G |
Masculine (stone):
| singular | plural | |
|---|---|---|
| nom | stān | stānas |
| acc | stān | stānas |
| gen | stānes | stāna |
| dat | stāne | stānum |
Neuter short (ship):
| singular | plural | |
|---|---|---|
| nom | scip | scipu |
| acc | scip | scipu |
| gen | scipes | scipa |
| dat | scipe | scipum |
Neuter long (word):
| singular | plural | |
|---|---|---|
| nom | word | word |
| acc | word | word |
| gen | wordes | worda |
| dat | worde | wordum |
Feminine short (gift):
| singular | plural | |
|---|---|---|
| nom | ġiefu [v] | ġiefa |
| acc | ġiefe | ġiefa |
| gen | ġiefe | ġiefa |
| dat | ġiefe | ġiefum |
A morph.xml file is a list of word-forms and a list of associated macros.
morph @name
> entry* @word @stem @pos @macros
> macro* @name
> fs? @id @attr @val
> lf?
The 'stem' attribute should be thought of as the semantic predicate (by default it's the same as the value of 'word'). The 'pos' attribute specifies which families are relevant in building the full lexical category. The 'id' attribute make reference to an identifier in the 'lexicon.xml' file.
| bed | bɛd | bed |
| blód | blo:d | blood |
| God | gɔd | God |
| land | land | land |
| nim | nɪm | take |
| ramm | ram | ram |
| be | bɛ | by |
| nan | nan | none/no |
| tíd | ti:d | time |
| in | ɪn | in |
| on | ɔn | on |
| up | ʊp | up |
| bin | bɪn | bin |
| swift | swɪft | swift |
| gold | gold | gold |
| grim | grɪm | grim |
| song | sɔŋg | song |
| corn | kɔrn | corn |
| lamb | lamb | lamb |
| wind | wɪnd | wind |
| þus | θʊs | thus |
| far | far | far |
| eft | ɛft | again/back |
| self | sɛlf | self |
| fyr | fyr | fire |
| sweord | swɛɔrd | sword |
| cræft | kræft | craft |
| léof | lɛɔf | dear (sir) |
| fréo | frɛɔ | free |
| hé | he: | he |
| him | hɪm | him |
| niht | nɪçt | night |
| hand | hand | hand |
| þurh | θʊrx | through |
| ic | ɪtʃ | I |
| cwæþ | kwæ&theta | said |
| cild | tʃɪld | child |
| gad | gad | goad |
| græg | græj | gray |
| cniht | knɪçt | boy/servant |
| betst | bɛtst | best |
| fisc | fɪʃ | fish |
| scip | ʃɪp | ship |
| ecg | ɛdʒ | edge |
| bricg | brɪdʒ | bridge |
| secg | sɛdʒ | man |
The following consonants are straightforward:
| p [p] | t [t] | k [k] |
| b [b] | d [d] | |
| m [m] | n [n] | |
| l [l] | r [r] | |
| x [ks] | ||
Fricatives each have two allophones:
| f [f/v] |
| s [s/z] |
| þ [θ/ð] |
| ð [θ/ð] |
The voiceless allophone appears initially, finally, and adjacent to a voiceless consonant. The voiced allophone appears between two vowels, or between a vowel and a voiced consonant.
The 'h' sounds has three allophones:
| h [h/ç/x] |
The [h] sound appears initially, [ç] after a front vowel, and [x] after a back vowel.
The sounds 'c' and 'g' have two allophones:
| c [tʃ/k] | g [j/g] |
The first allophone appears: (a) before a front vowel; and (b) at the end of a word, after a front vowel. The second allophone appears before a back vowel or a consonant.
Some special combinations:
| cn [kn] |
| gn [gn] |
| sc [ʃ] |
| cg [dʒ] |
Old English texts contain seven letters which denote vowels. There is evidence that each of these represents two distinct phonemes, one 'short' and one 'long':
| majuscule | minuscule | short | long |
|---|---|---|---|
| a | A | a or ɑ | ɑ: |
| e | E | ɛ or e | e: |
| i | I | ɪ or i | i: |
| o | O | ɔ or o | o: |
| u | U | ʊ or u | u: |
| y | Y | y | y: |
| æ | Æ | æ | æ: |
Note that there are two different proposed pronunciations for some of the short vowels.
The evidence for phonemic status of vowel length involves the existence of minimal pairs involving homographs like:
| written | short reflex | long reflex |
|---|---|---|
| ham | ham | home |
| is | is | ice |
| rod | rod | rood (i.e. cross) |
Because one reading of each homograph underwent a vowel change during the Great Vowel Shift, it is believed that these words were not homophones in Old English. In addition, we can use evidence of vowel shift in Modern English to identify long vowels in Old English.
In modern versions of Old English texts, long vowels are often marked with a macron: ā, ē, ī etc. The original Old English texts often use different glyphs: (a) 'i' without a dot; (b) 'y' with a dot; (c) a version of 'E' that looks a bit like 'C'; (d) sometimes a version of 'a' without the curly bit on top.
There are three digraphs/diphthongs, each of which can be either short or long:
| digraph | sound |
|---|---|
| eo | eo or eʊ |
| ea | æɑ |
| ie | ɪ |
Note that the digraph 'ie' was probably not pronounced as a diphthong. Long diphthongs/digraphs are conventionally designated by placing a macron above the first letter in the digraph.
Note also that the short vowels 'a', 'æ' and 'ea' are closely related - they evolved out of a single Proto-Germanic vowel 'a' (by phonological conditioning), and evolved back into a single Middle English vowel 'a'. This is important to know to understand certain inflectional paradigms, e.g. 'dæg' vs 'dagas'; 'geat' vs 'gatu'.
eel bee fee flee free knee beech beef been beep beer beet bleed bleep breech breed breeze cheek cheep cheer cheese deed deem deep deer feed feel freeze
A * is used when the pronunciations is not /i:/. Generally pronounced /ɛ:/ before the Great Vowel Shift, which rose to /e:/ then /i:/.
each ear earl* earn* earth* ease east eat eave flea beach bead beak beam bean bear* beard beast beat bleach bleak bleat breach bread* breadth* break* bream breast* breath* breathe cease dead* deaf* deal dean dear dearth* death* dread* dream fear feast feat freak gear gleam glean grease great* head* heal health* heap hear heard* hearse* heart* hearth* heat heath heave jeans leach lead[*] leaf league leak lean leap learn* lease least leave leaves mead meal mean means meant* meat
Others:
feather* heather* heaven* jealous* leather* leaven* meadow* measure*
The long/stressed vowels of Old/Middle English:
| front | back | |
|---|---|---|
| i: | u: | |
| e: | o: | |
| ɛ: | ɔ: | |
| a: |
Typical orthography for these was as follows:
| i: | y,i | mine,sight |
| u: | ou,u | house |
| e: | ee,e,e_e | sheep, me, mete |
| o: | oo,o | boot |
| ɛ: | ea,e | break, beak |
| ɔ: | oa | loaf |
| a: | aa,a | name |
The Great Vowel Shift involved everything moving up a slot:
macro @name > fs* > lf
private String name; private FeatureStructure[] featStrucs; private LF[] preds;
This class parses the entry elements in morph.xml files:
entry @word @pos (@stem) (@class) (@coart) (@macros) (@excluded)
There are the following fields:
private Word surfaceWord; private Word word; private Word coartIndexingWord; private String[] macros; private String[] excluded; private boolean coart = false;
The constructor is quite complicated. The value of 'macros' and 'excluded' is straightforward - you just split the value up into string tokens (macro names and excluded strings). The value of 'surfaceWord' comes from passing the value of the word attribute through the tokeniser in some strange way:
surfaceWord = Grammar.theGrammar.lexicon.tokenizer.parseToken(e.getAttributeValue("word"),coart);
The value of 'word' is derived as follows:
word = word.createFullWord(surfaceWord,
e.getAttributeValue("stem"),
e.getAttributeValue("pos"),
null,
e.getAttributeValue("class"));
This class parses member elements inside family elements inside lexicon.xml files:
member @stem (@pred)
The fields and constructor are obvious:
private String stem;
private String pred;
public DataItem(org.jdom.Element e) { ... }
If there is no 'pred' attribute, its value is the same as 'stem'.
This class is used to parse entry elements from inside family elements in lexicon.xml files:
entry @name (@stem) (@active) (@indexRel) > category
The fields for this class are:
private String name; private String stem; private Boolean active; private String indexRel; private Category cat; private Family family;
The constructor method is obvious too:
public EntriesItem(org.jdom.Element e, Family f) {
this.family = f;
...
cat = CatReader.getCat((Element) e.getChildren().get(0));
}
If there is no 'stem' attribute, this field gets the value Lexicon.DEFAULT_VAL, i.e. "[*DEFAULT*]". The default value of 'active' is Boolean.TRUE. If there is no 'indexRel' attribute, then the value of this field comes from that of the family. Obviously, the last line is the most interesting - it takes the first child only and converts it into a category? What if there are other children?
There are the usual 'get' methods, as well as a couple that get attributes of the immediate superfamily, as well a toString() method.
This class is used to parse XML family elements in OpenCCG lexicon.xml files:
family @name @pos (@closed) (@indexRel) (@coartRel) > entry* > member*
The fields of a Family object are thus as follows:
private String name; private String pos; private Boolean closed; private String indexRel; private String coartRel; private EntriesItem[] entries; private DataItem[] data; //members private String supertag;
The supertag is formed by removing slash modalities and other minor features from a category (i.e. from the family name). There are get and set accessor methods for each field. The constructor method is exactly as you'd think:
public Family(org.jdom.Element e) { ... }
See also: opennlp.ccg.lexicon.EntriesItem and opennlp.ccg.lexicon.DataItem.
The constructor method of the opennlp.ccg.grammar.Grammar class contains the following code:
public final Lexicon lexicon; lexicon = new Lexicon(this); lexicon.init(lexiconUrl,morphUrl);
Here we turn to the code for the opennlp.ccg.lexicon.Lexicon class. Here are two basic fields and the constructor:
public final Grammar grammar;
public final Tokenizer tokenizer;
public Lexicon(Grammar g) {
this.grammar = g;
this.tokenizer = new DefaultTokenizer();
}
The init method starts as follows, where there are two input parameters, lexiconUrl and morphUrl:
List<Family> lexicon = null; List<MorphItem> morph = null; List<MacroItem> macroModel = null; lexicon = getLexicon(lexiconUrl); Pair<List<MorphItem>,List<MacroItem>> morphInfo = getMorph(morphUrl); morph = morphInfo.a; macroModel = morphInfo.b;
See also opennlp.ccg.lexicon.Family, opennlp.ccg.lexicon.MorphItem, and opennlp.ccg.lexicon.MacroItem.
This class implements the notion of a (multiple inheritance) hierachy of syntactic types. It is called by the constructor method of the opennlp.ccg.grammar.Grammar class, as documented here. Each Grammar object has an instance field types of class Types. The constructor method for Grammar calls one of two constructor methods to initialise this field: (a) new Types(URL u,Grammar g); or (b) new Types(Grammar g).
The code for the
Apart from that, the code is a bit of a mess, so I'm going to ignore these classes for the moment.
This class encodes the notion of a CCG grammar, essentially a lexicon, a set of rules and a type hierarchy. Its constructor is called by the opennlp.ccg.TextCCG class, as discussed here. Here are the essentials (ignoring XSLT files for transforming from and to LFs, pitch accents, special tokenisers, supertags etc.):
public final Types types; public final Lexicon lexicon; public final RuleGroup rules; private String grammarName = null; public static Grammar theGrammar; //nasty hack!Here are the basics of the constructor, which takes a single input parameter url, of class java.net.URL:
theGrammar = this;
SAXBuilder builder = new SAXBuilder();
Document doc = builder.build(url);
Element root = doc.getRootElement();
grammarName = root.getAttributeValue("name");
...
Element typesElt = root.getChild("types");
URL typesUrl;
if (typesElt!=null)
typesUrl = new URL(url,typesElt.getAttributeValue("file"));
else typesUrl = null;
Element lexiconElt = root.getChild("lexicon");
URL lexiconUrl = new URL(url,lexiconElt.getAttributeValue("file"));
Element morphElt = root.getChild("morphology");
URL morphUrl = new URL(url,morphElt.getAttributeValue("file"));
Element rulesElt = root.getChild("rules");
URL rulesUrl = new URL(url,rulesElt.getAttributeValue("file"));
...
if (typesUrl!=null) types = new Types(typesUrl,this);
else types = new Types(this);
lexicon = new Lexicon(this);
lexicon.init(lexiconUrl,morphUrl); *****
rules = new RuleGroup(rulesUrl,this);
...