1. Introduction
Plaid UD is a web-based editor for building Universal Dependencies treebanks. This guide is written for annotators who will use the editor to do annotation work. It walks through the everyday loop (create or import a document, tokenize it, annotate it, and export it as CoNLL-U), and then points you to the more advanced features.
Plaid UD is one of the example apps built on the Plaid platform. You don’t need to know anything about Plaid itself to use the editor, but if you’re curious about the underlying platform, or you want to script against your data, see the Plaid manual and the Python / JavaScript client references.
|
Note
|
Plaid UD is a demonstration app. It covers the core treebank-annotation workflow well, but it deliberately leaves some UD features out of scope (for example empty nodes and the MISC column). Those are noted where they come up. |
2. Getting Started
Plaid UD is served by your Plaid server. There is nothing to install.
Open the editor in your browser at the /ud/ path on that server.
If you are running Plaid locally with the default settings, that is:
http://localhost:8080/ud/
Sign in with the email address and password your project administrator gave you. Your session stays signed in until the token expires or you sign out. Your name at the top right opens your account, with your Profile and Sign out.
The editor has three levels you move between:
-
Projects: the list of all projects you can access (your home screen after signing in).
-
Documents: the texts inside one project.
-
A document’s own tabs:
Text Editor,Annotate,Export,CommentsandDetails, switched at the top of an open document.
Project-wide tabs sit alongside the document list inside each project: Search, Guidelines, Assistant, Validation, Activity, Settings and Import and export. Validation, Activity and Settings are for maintainers, and Assistant appears only when one is running. Anyone else who opens Validation, Activity, General or UD settings by its address is sent back to the project list with Only a project’s maintainers can open this page. The assistant is also a panel you can open beside any screen (see The panel).
3. Projects
A project in Plaid UD holds a collection of documents that share one annotation scheme. Universal Dependencies describes a text at three levels, and the editor follows the same model:
| Level | What it is |
|---|---|
Sentences |
Spans that tile the whole text, one per sentence. |
Tokens |
The orthographic words you see in the text, the UD tokens. A token may be a multi-word token (MWT), e.g. Spanish del standing for de + el. |
Words |
The UD syntactic words, the numbered rows of a CoNLL-U file. All annotation lives here: part-of-speech, lemma, features, and the dependency relation. An ordinary token is a single word. A multi-word token contains several. |
For most tokens there is exactly one word, so the two coincide and you can ignore the distinction. The split only matters for multi-word tokens, where one written token expands into several annotated words.
3.1. Creating and Opening Projects
From the Projects screen you can sort and search the list, and create a new project with New UD project. A new project is ready for Universal Dependencies annotation right away. Click it to open its documents.
4. Documents
Open a project to see its documents. As with projects, you can sort and search the list and create new documents. A new document starts empty. You give it a name, then add text in the Text Editor (next section) or by importing.
4.1. Importing CoNLL-U
If you already have annotated data, you can import it instead of typing it in.
Open the project’s Import and export tab and drop .conllu files on it, one or many, or a ZIP of them.
Each # newdoc boundary inside a file becomes its own document.
The same tab exports the whole project as a ZIP of .conllu files.
|
Warning
|
Import brings in tokens, sentence metadata, and the dependency tree, but two CoNLL-U features are not stored and will be dropped on import: empty nodes and the MISC column. If you need those preserved, Plaid UD is not the right tool for that file. |
The DEPS column is read into the enhanced dependency graph (see Enhanced Dependencies).
An enhanced dependency whose head is an empty node is dropped along with the node.
A word whose only enhanced heads were empty nodes keeps the relation its tree gives it, which for a gapped clause is orphan.
A word whose DEPS is empty, in a sentence where other words have one, has no head in the enhanced graph: its tree relation is one the graph leaves out. In a sentence with no DEPS at all, the file simply says nothing about the enhanced graph and every word follows its tree.
5. Tokenizing Text (Text Editor)
The Text Editor tab is where raw text becomes tokens.
5.1. Entering and Tokenizing Text
Paste or type your text into the editor, one sentence per line. Use Save to save edits to the raw text without re-tokenizing. Each change is saved where you typed it.
When someone else saved the text after you started editing it, your save puts your changes onto their text.
When you both changed the same passage, the save is refused and the editor says Changed elsewhere in the same passage. Discard changes then shows the saved text, and you make the edit again.
The save is also refused when the text has passages that read the same and your change could be in either of them.
Tokens already made follow an edit of the saved text. A change inside a token, or touching it with no space between, changes that token, and it keeps its words and annotations. Typing a space inside a token splits it there: the token and its annotations stay on the part holding most of its letters, and the rest is untokenized text. Spaces a token already had do not split it. On a project whose tokens are shared with another app, the split holds for that app too. Deleting the space between two tokens keeps them two tokens. Merge them to make one. A token goes, along with its annotations, only when all its text is deleted. Text typed right before a sentence’s first character joins that sentence, unless it holds a line break, and text typed right after a sentence’s last character joins that one.
There are two ways to turn saved text into tokens, and both are runs (see Services):
-
Tokenize segments the text into sentences and tokens. The built-in method uses your project’s tokenizer locale, falling back to the project’s language (both set in project settings, and language-agnostic when neither is set). It makes one sentence of each line of the text. Each word gets a lemma copied from its form, marked as machine-made, which Parse replaces. You add the other annotations and the dependency tree yourself in the Annotate tab.
-
Parse does the whole job at once. Where your project has a parsing service connected, for example the bundled Stanza parser, it fills in the entire document: sentences, tokens, words, every annotation, and the dependency tree. This gives you a first draft to correct by hand, and is the fastest way to get started.
Each button opens a dialog naming the method it will use and the options that method offers. A document that already has tokens cannot be tokenized again. Tokenize says to clear the tokens first, and Clear tokens beside it does that, after asking. Clearing takes the annotations on those tokens with them, and cannot be undone.
5.2. Working with Tokens
After tokenizing, the text is shown with each token as a clickable chip:
-
A green border marks the first token of each sentence. Click any token to toggle whether a sentence starts there. This is how you fix sentence boundaries. Splitting a sentence removes the dependency relations that would cross the new boundary, wherever the split is made.
-
A teal chip marks a multi-word token (a token that contains more than one word).
Hover over a token to open its panel, where you can:
-
toggle Start of sentence
-
edit the token’s words, splitting it into several to make a multi-word token or merging them back into one
-
delete the token
You can also select a span of text directly to create a new token from it.
Editing a token’s words with the same number of words keeps their annotations and relations, and changes only the forms. Changing the number of words replaces them, and asks first when annotations would go. Deleting a token asks first when anything on its words would go, including the split of a multi-word token or a word’s changed form.
6. Annotation
The Annotate tab is where the real work happens. It shows each sentence as a row of editable cells above an interactive dependency tree.
Sentences come 25 to a page, with a pager above and below them. A document reopens on the page you left it on, and anything that reaches a particular sentence (a search result, a review sweep, a link from the assistant) turns to its page first. Sentence numbers belong to the document, not the page, so sentence 40 is sentence 40 wherever it is shown.
6.1. Tag and Lemma Editing
Each word has cells for:
| Column | Meaning |
|---|---|
UPOS |
Universal part-of-speech tag. |
XPOS |
Language-specific part-of-speech tag. |
Lemma |
The word’s dictionary form. |
Features |
Morphological features, as |
Click a cell to edit it. The UPOS, XPOS and Features cells, and the dependency-relation labels in the tree, offer autocomplete from your project’s controlled vocabularies. By default those lists are suggestions and a value off the list is accepted, but a maintainer can close any one of them so only listed values are (see Suggestions or rules).
Move around the grid with Tab, Enter, Escape and the arrow keys. Arrow Up out of the top row jumps into the dependency tree. The full list is in Keyboard Shortcuts.
6.2. Tree Editing
Below the grid, each sentence’s dependency tree is drawn as arcs between words.
-
Draw an arc: press and drag from one word to another. The word you start on becomes the head, and the word you release on becomes its dependent.
-
Mark the root: drag from the ROOT bar to a word.
-
Rename a relation: click an arc’s label to edit its dependency relation (
nsubj,obj, …), again with autocomplete. -
Delete a relation: click its label, then the bin to the left of the box that opens, or press Shift+Backspace before typing anything.
-
Keyboard: from the grid, Ctrl/Cmd+D jumps focus into the tree’s labels. Tab and Shift+Tab walk between them, and from inside an open label they commit it and open the next. See Keyboard Shortcuts for the rest.
6.3. Enhanced Dependencies
Every sentence has an enhanced dependency graph beside its tree. The enhanced graph starts out equal to the tree, and you record only where it differs. A project that never does so exports a DEPS column that restates the tree.
-
Add a relation: hold Ctrl or Cmd as you start to draw an arc. It is drawn under the words as you go, and you can let go of it on a word or anywhere beneath one. The word you release on keeps its head in the tree and gains a second one in the enhanced graph. This is how a subject shared by two coordinated verbs, or the subject of a controlled verb, is recorded. The tree is drawn above the words and these relations below them. A label there is edited and deleted like any other.
-
Give a relation a different label in the enhanced graph: start with Ctrl or Cmd held and draw over a relation the tree already has. Its label opens, and what you type becomes the label in the enhanced graph, such as
nmod:offornmod. The tree keeps its own relation, which is shown faded. Deleting that label puts the tree’s relation back into the enhanced graph. -
Leave a relation out of the enhanced graph: Ctrl/Cmd+click a relation of the tree, or press Ctrl/Cmd+E on its label. The relation is shown faded. Do the same again to put it back.
A change to the tree carries over to the enhanced graph by itself, because only the differences are stored. Splitting a sentence removes the enhanced relations that would cross the new boundary, as it does the tree’s.
On export, DEPS lists every relation of the enhanced graph for each word.
A project made before enhanced dependencies existed gains them the first time one of its maintainers opens a document or runs an import.
Until then its editor has no enhanced gestures, and an import into it drops DEPS with a notice.
Empty nodes are not supported, so an elided predicate cannot be added, and gapping stays as the tree’s orphan relation.
6.4. Metadata
A treebank records more than annotations. CoNLL-U keeps those notes as comment lines, and Plaid UD edits them in two places.
About a document, on its Details tab: where the text came from, its genre, its licence, whatever the project decides to keep.
About a sentence, under the sentence in the Annotate view, behind Edit metadata in the row of actions below the grid. Every sentence carries sent_id, CoNLL-U’s own identifier for it, whether the project declares anything else or not. A free translation is conventionally text_en, naming the language it is in.
Which fields a project offers is a maintainer’s choice, under Settings → UD settings. A field name cannot contain a dot, and a sentence field cannot be called sent_id or text, which the exporter writes itself.
Fields save as you leave them, and clearing one removes the note rather than storing a blank. A value already stored under a name the project no longer offers is still shown, marked, so it can be read and cleared instead of quietly exporting forever.
Everything here goes out with the document: # source = … above a document, # sent_id and # text_en above a sentence.
6.5. Text Direction
A document’s Details tab also holds Text direction, which says which way its tokens are laid out.
It is set to Automatic, which reads the document’s own text. A treebank in Arabic, Hebrew, Thaana, N’Ko or any other right-to-left script lays its tokens out right to left in the annotation grid, with the row labels on the right and the dependency arcs over the words where they now are. Set it by hand for a document the text cannot speak for, such as a right-to-left language written here in a Latin transliteration.
Direction affects the way tokens are laid out, not the values in them. Every field decides for itself, so a UPOS tag or an English lemma reads left to right in a column standing under an Arabic word. Arrow keys follow the words: in a right-to-left sentence the next token is the one to the left, so it is the left arrow that moves forward through the sentence.
6.6. What This Project Has Said Before
In a Lemma, XPOS or Features cell, Alt+Down lists the values this project has already used for a word like this one, with how many times each.
The lemma cell asks about the form, meaning what the project has called this string before. The other two ask about the lemma, because once that is settled the tag and the features usually follow, and the same string can be two lemmas.
Pick one with the arrows and Enter, or click it. Escape goes back to the cell. A word with nothing to compare against simply shows nothing.
This is consistency help for hand annotation. A connected parser writes a first draft of most of this already.
6.7. Moving Between the Text and the Annotations
Tokenizing and annotating are one job done in two places, so each hands over to the other at the sentence you are looking at.
From Annotate: Edit text under a sentence, or Alt+click any word in it, opens the Text Editor scrolled to that sentence.
From the Text Editor: Alt+click any token opens Annotate at that token’s sentence. A plain click still toggles the sentence boundary.
Either way the sentence you arrived at is briefly outlined, so you can see where you landed.
6.8. Machine vs. Human Annotations
Annotations nobody has checked yet are marked, so you can see at a glance what is left to review. There are two marks, and they say where the annotation came from:
-
violet, dashed: an automatic parser produced it.
-
amber, dashed: a member whose work the project reviews (see Sharing and Permissions) entered it. This mark appears only in projects that review somebody.
Anything a trusted member has entered or confirmed is shown plainly, in its own colour. Hovering a cell or a relation names its state and its producer. Any edit you make to a marked cell or relation confirms it and clears the mark. If your own work is reviewed, the mark stays on what you enter until a maintainer or another trusted member confirms it, and accepting a prediction records it as your contribution rather than confirming it.
Typing a value again is an edit, even where the value does not change.
In a Features cell that means writing the same Key=Value pair, which confirms the feature rather than adding a second one.
Passing through a cell with Tab or an arrow key, and leaving one with Escape, confirm nothing.
If a prediction is already correct, you don’t have to retype it:
-
Ctrl+Enter (or Cmd+Enter) accepts everything proposed on the word and moves to the next word, so a sentence can be reviewed with one hand.
-
Ctrl+Backspace throws the parser’s work on the word away and moves on, for a proposal that is wrong wholesale rather than worth correcting cell by cell. Work entered by a person is never deleted this way.
-
Ctrl+Shift+↓ and Ctrl+Shift+↑ jump to the next and previous word that still needs a look, across sentences.
-
Accept predictions and Discard predictions, under each sentence, do the same for the whole sentence.
Gestures and marks, above the first sentence, opens a legend of these marks and every gesture in the editor.
6.9. Saving
A cell is saved when you leave it, and a relation when you draw or change it. There is no save button. A save that gets no answer, because the connection is down or the server does not reply, is sent again until the server answers, and the edits made after it wait their turn. Meanwhile the page reads Offline, retrying, or Can’t reach the server, retrying while the browser is online, and closing the tab asks first. A save that reached the server before its answer was lost is not saved twice.
Two people can have one document open, and neither page shows the other’s edits as they are made. An edit is refused when someone else changed the same kind of annotation in the document since the page loaded (a lemma, a head, the words), even on another word. A change to another kind of annotation, on other words, does not refuse it.
When someone else changed the same cell first, the cell shows their value, with yours under it: Yours: X · Enter to keep yours.
A message names the change, such as b changed this to NOUN.
Enter in the cell saves yours, Escape or typing another value drops it, and leaving the cell saves nothing.
The same happens when the word was split or joined by someone else before your value was saved.
When the word is gone by the time your value is refused, a message says Not saved: and the value.
A value refused because of a change elsewhere in the document stays in its cell, and leaving the cell again sends it.
The message reads Changed elsewhere. Your value is in its cell, not saved.
Any other edit that is refused says Changed elsewhere. Now showing the latest version. Redo your edit., and the document shows the other person’s change.
A repair is made when a document that needs one is opened.
When someone saves while it is being made, it starts again by itself.
If it fails again, or the server does not answer, a message asks you to reload the page.
A document with no sentences to show until then says The document could not be repaired. Reload the page to try again.
7. Export
The Export tab shows the current document as CoNLL-U, with Copy and Download.
To export a whole project, use its Import and export tab, which packages every document into one ZIP of .conllu files.
8. Guidelines
The Guidelines tab is the project’s own annotation manual: the conventions the people working on it have agreed to, written down where everyone can read them. The controlled value lists under UD settings say which tags a column takes. Guidelines are for what a list of tags cannot say: which of two analyses this treebank uses for a construction, what belongs in the lemma column, how a hard case was decided and why.
Everyone who can open the project can read them, and anyone who can annotate can write them.
Each guideline has a title and a body written in a rich-text editor (bold, headings, lists, tables, links). The title is the only thing to keep current, and it is how a guideline is referred to, including by the Assistant. Name the subject rather than the rule, so that the title still fits after the rule is refined. Two guidelines with the same title are worth avoiding. The editor says so while you are typing one that is already in use, and saves it anyway. If a title cannot name what a guideline covers, it is two guidelines. One topic each is what makes them findable, by a person and by the Assistant.
Open the body with the rule itself. When the manual grows too large to send to the Assistant whole, that first line is what it sees beside the title and chooses on.
The editor writes Markdown, and Markdown switches to the source if you would rather type it.
Pinning a guideline sends it to the Assistant in full on every question. Unpinned guidelines are sent in full too while the manual is small, and the Assistant opens them by name once it is large. Under each of its replies is a line saying how much of the manual that answer had in front of it.
If somebody else saves a guideline while you have it open, your save is refused rather than overwriting theirs. Nothing you have typed is lost: the editor keeps it, and saving a second time overwrites theirs deliberately.
Guidelines are not part of a project export here, which is a ZIP of CoNLL-U files and has nowhere to put them. A project shared with plaid-igt, whose archive does carry them, keeps them that way.
The Assistant can draft a guideline as well as read one. If you tell it a convention while asking about something else, it will propose writing it down, and you approve or discard it like any other change it proposes. It will not do this for a decision about a single word, or for something it worked out from the data on its own. When it changes a guideline that already exists, the card shows you the exact passage it is replacing and what with. A change that replaces the whole guideline is marked Rewrite instead, because there the previous wording is not on the card for you to compare.
Every change to a guideline is recorded in the project’s Activity, with who made it and when, including the ones the Assistant drafted.
9. Assistant
The Assistant tab is a chat with a language model that knows your project. Ask it which words still have no lemma, whether any dependency relations look wrong, or to summarize the parts of speech across the corpus. It reads the documents, the search results, the project’s guidelines and the history to answer, and it shows what it read.
When you ask it to change something, it does not touch the data. It proposes a plan, which you Approve or Discard. Every change is listed under the document it lands in, naming the word and its reference, and each one links to that sentence in the annotation editor, so a plan can be checked before it is approved. A change to something a person made or accepted is marked Accepted, counted on a line of its own above the list, and never folded away. A replacement across the corpus counts the values it replaces that a person made or accepted, and says so on its row. Approved changes are written under your own account, recorded as the assistant’s work accepted by you, and show as accepted. Tick Record as human-made before approving to record them as your own work instead. Reading needs Reader access, and applying a plan needs Writer.
Approving a plan after a sentence it changes has been edited is refused, and the card then reads Out of date. An edit to another sentence does not stop it. A replacement across documents is also out of date when a document starts or stops matching it after the plan was made. Approving a plan on a document that a parser run is writing to is refused with nothing written. Approve again once the run has finished.
A plan that stops partway reads Partly applied, with how many of its changes were written, and the changes not written are dimmed. Ask the assistant to finish it. A parse the assistant started that stops partway reads Partly applied too, with the number of sentences parsed. Ask the assistant to parse again to finish.
A reply lands whether or not the page stays open: leave the tab, reload, or close the browser, and the answer is there when the conversation is next opened. Stop ends a turn early, even while the model is still working on its answer. Conversations are private to you, kept per project, and can be downloaded or copied as Markdown. Your server’s administrator can read them.
9.1. Examples in a reply
When a reply rests on a particular sentence, that sentence is shown, and it can be drawn three ways. Tree gives the dependency arcs over the words, with the relation on each arc and an arrowhead on the word it points at, as the UD documentation draws them. Grid gives only the columns the point is about. Table gives every column the sentence fills.
The assistant picks whichever fits what it is saying, and you can switch. A sentence nobody has parsed is not offered a tree. The heading of each example opens that sentence in the annotation editor.
9.2. The panel
The Assistant button in the header opens the same assistant in a panel beside whatever you are looking at. It is there wherever there is a project for it to be about, and on the general screens, where it asks which project to use. The screens for making a project and for importing one have neither, so it is not offered there. A handle on the right edge of the window, halfway down, opens it too: it widens under the pointer to show the assistant’s mark. The panel stays where it is as you move around, so a question can be asked without leaving what you are reading and an answer can be read while you get on with something else. The page narrows to make room for it. The panel can be dragged wider or narrower and hidden again, and its width and whether it is open are remembered. It starts closed until you first open it. A narrow window has no room for it, and neither the button, the handle nor the panel is offered there.
One conversation runs per project. The panel picks up the last one you had in the project you are in and keeps it as you move from document to document. Each question records where it was asked from, so a question about "this sentence" still means the right document when the conversation is read again later.
The panel’s history lists the conversations you have had in the project, with All projects to widen the list past it. Opening one from there keeps you on the screen you are on. ⤢, beside it, moves the open conversation to the Assistant tab and closes the panel.
A question asked on a document is about that document: name no document and that is the one it reads, though it can still look at the rest of the project when the question calls for it, to compare or to count. Ask, beside Edit text under a sentence, puts that sentence into the question, and you can take it out again before sending.
Typing @ in the message box names something without leaving the message: the sentences of the document you have open, and the documents of the project. Sentences are matched on what they say, so @ followed by a word finds the sentence that contains it. Arrow keys move through the list, Enter takes the one that is highlighted, and Escape closes the list and leaves what you typed. An example that names a sentence of the open document scrolls the annotation editor to it rather than opening a new tab.
On a screen with no project of its own, the panel keeps the conversation it has. Before you have opened a project at all it asks which one to use, because the assistant reads one project at a time.
A plan approved in the panel re-reads the document, so the annotation shows the result without losing your place.
9.3. When no assistant is connected
The assistant is a service that your server’s operator runs and picks the model for (plaid-ud-agent, installed from a checkout of the source tree).
When none is connected there is nothing to open, so the tab, the button, the handle and Ask are not shown.
Conversations already saved are unaffected and open from their own links.
10. Additional Features
10.1. Search
Each project has a Search screen that finds structures across all its documents using Grew-style patterns. For example, this finds verbs with a nominal subject:
pattern { V [upos=VERB]; S [upos=NOUN|PROPN]; V -[nsubj]-> S }
Type a pattern and run it with the Search button or Ctrl/Cmd+Enter. Results are grouped by document. Click a match to jump straight to that sentence in the Annotate view. A built-in syntax reference with more examples is available on the search screen.
10.1.1. Searching the enhanced graph
A relation that the enhanced graph adds to the tree carries Grew’s E: prefix: E:nsubj is an enhanced nsubj.
-
V -[nsubj]→ Sreads the tree alone. -
V -[E:nsubj]→ Sreads only the relations added in the enhanced graph. -
V → S,V -[1=nsubj]→ SandV -[^det]→ Sread both. Addenhanced=yesor!enhancedto a feature label to read one. -
V -[nsubj|E:nsubj]→ Sreads each label where it belongs.
This is how Grew itself reads a CoNLL-U file: a DEPS entry that repeats the tree is the tree’s relation, and only the others are marked enhanced.
A relation of the tree that the enhanced graph leaves out or relabels is still a relation of the tree, so a pattern finds it under its plain label.
is_tree, is_projective and an unlabelled X →> Y are about the tree and ignore the added relations.
A rule writes with the same labels (see Changing many sentences at once): add_edge V2 -[E:nsubj]→ S adds a relation to the enhanced graph, and e.enhanced = yes on a named edge moves it there.
Added between two words the tree already joins, it replaces the tree’s relation in the enhanced graph, exactly as drawing it in the editor does, and the preview says which relation is left out.
10.1.2. Looking for a word
Structure needs Grew, and a word does not. The box above the pattern looks in a single field (form, lemma, UPOS, XPOS, dependency relation or features) for text that contains, is, or matches (as a regular expression) what you type.
It does not run a search of its own: it writes the Grew pattern and hands it to the box below. So a quick lookup is the first draft of a real one. Run it, then edit the pattern it wrote.
10.1.3. Counting instead of reading
After a search, Count by re-runs the same pattern as a tally: pick a pattern node and a field (V.upos, S.lemma, or a named edge’s e.label) and the result is a sortable table of values with how many matches each accounts for.
The matches themselves are never fetched for this, because the server does the counting, so it stays fast on a corpus too large to page through.
Clicking a value narrows the search to it. The clause goes into the pattern in the box, where you can see it and edit it, and the search re-runs. A named edge’s label is the exception: it belongs in the arc rather than on a node, so those rows are not clickable.
10.1.4. Changing many sentences at once
A pattern followed by a commands block is a rule, and the Search screen previews what it would change instead of listing matches.
This one retags that wherever it was tagged as a subordinator:
pattern { X [lemma="that", upos=SCONJ] }
commands { X.upos = PRON }
With a rule in the box the button reads Preview changes. A preview writes nothing.
It lists every sentence the rule would change, grouped by document, with one line per change such as that: upos SCONJ → PRON.
A sentence the rule applied to more than once shows the number of times. Click a sentence to open it in the Annotate view.
Tick the sentences or whole documents you want, or use Select all, then press Apply. Applying is for maintainers. Anyone who can search can preview.
A rule is applied to a sentence again and again until it no longer matches, as in Grew.
So a rule has to stop matching what it has just changed. The rule above does, because the word is no longer SCONJ. A rule that adds something usually needs a without block that names what it adds:
pattern { V1 -[conj]-> V2; V1 -[nsubj]-> S }
without { V2 -[E:nsubj]-> S }
commands { add_edge V2 -[E:nsubj]-> S }
A rule that matches and changes nothing, or that is still matching after 1000 applications in one sentence, stops with an error on that sentence. A sentence with an error cannot be selected, and the rest of the preview is unaffected.
The commands cover features (X.upos = VERB, X.Number = D.Number, del_feat X.Number), relations (add_edge, del_edge, e.label = "obj", shift), and deleting a word (del_node). Words cannot be added or reordered, since they come from the text.
Several rules can be written as rule name { … } blocks, which are tried in order unless a strat main { … } says otherwise. The syntax reference on the Search screen lists every command with examples.
Some things to know before applying:
-
The changes count as yours. A rule that changes a value a parser wrote confirms it, the same as editing it by hand.
-
Grew does not require a tree, and neither does a rule. A relation added to a word that already has a head gives it a second one, and the preview warns that the word has 2 heads.
-
Where a vocabulary is closed (see Suggestions or rules), a sentence in which the rule writes a value off the list is refused, and the preview shows which value.
-
Each document is changed in one step, recorded in its history as Rewrite: followed by the rule names. To undo a rewrite, restore the document to the entry before it (see Restoring an earlier state).
-
If someone else changes a document between your preview and your apply, that document is left untouched and the run stops there. Documents already done stay changed. The preview then runs again, so you can apply the rest.
After an apply the preview runs again by itself. No sentences to change means the rule has nothing left to do.
10.2. History
Every change is recorded.
History at the end of a document’s tab row, on any tab, opens its history, where you can select an earlier point to view the document exactly as it was then.
Each entry is one action as you performed it (for example "Update upos", "Create relation", or a parser run), even when it was made up of several underlying changes.
A parser run is listed under the service’s name and names the person who ran it: Stanza UD parse (en), requested by Ana.
Selecting an entry shows the document as it was right after that action. Entries made of several changes have an arrow you can expand to see, and select, each individual change.
Historical views are read-only, on every tab that can show them: Annotate, Export and Details.
The Text Editor and Comments are not shown at a past state.
Return to current goes back to the live document.
An entry that is one of these kinds is marked with it:
-
Import: a file brought in.
-
Automatic: a parser’s run, or another service’s.
-
Assistant: an approved assistant plan.
-
Bulk edit: a Grew rewrite (see Changing many sentences at once).
-
Review: a prediction accepted as it stands.
-
Repair: a repair made when the document was opened.
The Activity tab marks its entries the same way.
An assistant plan that stopped partway is listed as Assistant, partly applied: followed by the changes it wrote.
10.2.1. Restoring an earlier state
Maintainers can put a document back the way it was. While viewing a historical state, choose Restore. Before anything is written, the dialog lists what would change: sentences, tokens, words, annotations in each field, and dependency relations.
Restoring is one entry in the history, so the state from just before it stays available. The toast that confirms a restore also offers Undo for fifteen seconds.
Whatever was deleted since the chosen moment comes back under its original identity, so comments attached to it stay attached. Two things cannot come back: annotations whose layer has since been deleted, and links to vocabulary entries that no longer exist. The dialog names these before you commit, and the toast repeats them afterwards.
10.3. Sharing and Permissions
Projects are shared by adding members with a role:
-
Reader: can view everything but make no edits.
-
Writer: can edit documents and annotations.
-
Maintainer: can also change project settings, manage members, and delete the project.
Add members from the project’s Access settings by searching for a name or email address. When you only have Reader access, the editor shows everything but its controls are inert.
Each member also has a Review work checkbox, independent of their role. Tick it and everything that person enters is marked as needing review until someone else confirms it (see Machine vs. Human Annotations). The mark is shared with every Plaid app on the project.
10.4. Settings
A project’s maintainers can tailor the annotation scheme from the project’s Settings tab:
-
UD settings: the controlled vocabularies (UPOS, XPOS, DEPREL tags), tag and relation colors, the feature inventory offered in the Features column, and the document and sentence fields the project keeps notes in (see Metadata). All vocabularies start from the standard universal sets and can be extended or replaced. Save writes only the settings you changed. If another maintainer changed the same one since you opened the tab, the save is refused and the tab shows the stored settings.
-
General: the project’s language, the tokenizer locale used by Tokenize, and project deletion. The language is a BCP-47 tag such as
en,deorzh-Hans. It is the language a connected parser starts on, and the locale Tokenize segments with unless the tokenizer locale is set to something more specific.
10.4.1. Suggestions or rules
A vocabulary offers its values and accepts anything else. That is the default, and for most projects it is the right one: a value nobody expected is usually worth seeing rather than refusing.
Where a project wants the list to be binding, the switch beneath it makes it so. It reads Refuse tags outside this list for UPOS and XPOS, Refuse relations outside this list for dependency relations, and Refuse features outside this list for the feature inventory. What refusing means differs by field:
-
UPOS and XPOS: the tag must be on the list.
-
Dependency relations: the BASE relation must be on the list. A subtype is judged by its base, so
nsubj:passis allowed wherevernsubjis, and a project does not have to list every subtype its language uses. -
Features: the key must be in the inventory and the value in that key’s list. A key listed with no values accepts any value.
A closed list is enforced where a person types: the annotation cells and the Grew rewrite preview. A closed UPOS, XPOS or relation list is also enforced by the server, on every write. A value off the list is refused, even from a page opened before the list was closed, and the cell shows the stored value again with a message naming the value. An approved assistant plan and a script are refused the same way. Three kinds of value are kept anyway: values brought in by an import, values a document copy or a restore writes back, and a parser’s predictions nobody has accepted yet. Accepting such a value is refused while it is off the list. Off-list machine output is a signal, and refusing it at the door would hide it. A closed feature inventory is enforced only where a person types. Spaces at either end of a value, no-break spaces included, are ignored when it is checked.
Closing a list, or removing a value from a closed one, is refused while any other value in the project’s annotations is off it, and the message names one of them. Change them first, or add them to the list.
10.4.2. What a value is for
Each value can carry a one-line definition, shown beside it while picking one. The 17 UPOS tags and the 37 universal relations ship with theirs. Edit any of them, or describe your own values, under the list. A definition you change is stored with the project, so a value you never touched follows the app.
Features are described on the whole pair, Number=Plur rather than Plur, since that is what the picker offers and what a word stores. A key with no values listed accepts anything, so it has no pairs to describe.
10.5. Comments
A document has a Comments tab: every thread on it, the document’s own first, then one per sentence.
Comments are for talking about the work, not annotating it. They are not part of the data: they never change a document, they do not appear in an export, and they are not in its history.
Where they can be attached is deliberately limited to a sentence and the document. There are no per-word threads, and a note about one word belongs in the sentence’s thread where the context is.
From the Annotate view, each sentence carries a small badge under its grid: the number of comments on it, or an invitation to write the first. It opens the thread in place, so a remark can be made without leaving the text.
A thread is collapsed to its latest comment, and opens on click. Bodies are Markdown. You may edit and delete your own, and a maintainer may delete anyone’s. Readers can read comments but not write them.
A comment outlives what it is about. If the sentence it hangs on is merged away or re-segmented, the comment stays, marked Outdated, under the caption it was posted with. Those are shown separately from the current ones.
While a thread is open, comments from other people appear as they are written.
10.6. Validation
Maintainers get a Validation tab: every value the project has stored that its own vocabularies do not list, with how many times it occurs.
It exists because an open list takes any value, and a closed one still keeps values from an import, a copy or a restore, and a parser’s predictions nobody has accepted (see Suggestions or rules). A closed feature inventory holds only where a person types. This is where you see what got in.
Each value opens to the sentences it occurs in, and each of those is a link straight to that sentence in the editor. Features open the same way, on the whole Key=Value pair.
Where the list is simply incomplete rather than the values wrong, Add N values puts them on it. That is offered, never done for you: a project may well want the list as it is and the values changed instead.
Two things it deliberately does not report: a relation subtype whose base is listed, since nsubj:pass is already legal wherever nsubj is, and a feature value under a key listed with no values, since such a key accepts any value.
10.7. Activity
A project’s maintainers get an Activity tab: who has been working on it, and on what.
It shows a tally per person over a window (7, 30 or 90 days, or all time) with the changes they made, the documents they touched, and when they were last seen, alongside the members who have not appeared in that window at all. Below it is the project’s own history, newest first, with each entry linking to the document it changed.
10.8. Administration
The server administers itself through one area, and it lives in Plaid IGT: the release bundles both apps together, so a second one here would be a second answer to the same questions. Administrators reach it from Admin in the header.
It covers accounts (creating people, deactivating them, issuing password-reset links), invitations, every project on the server and who is on each, the services connected, and server-wide activity. A project set up for Universal Dependencies is labelled UD in its project list, and its name opens it back here.
10.9. Services
Plaid UD can call external services for certain tasks. A service is a small program, often a Python script wrapping a model, that connects to your project when it starts and advertises which task it handles. Once connected, it appears wherever its task is offered, and a maintainer can make it the default for that task under Settings → Services. Some tasks ship with a built-in method that needs no service at all.
Both of them work the same way. A button names the run, and it opens a dialog with the Method to use, that method’s options, and the button that starts the run. A run keeps going if you close the dialog, and the button that opened it shows how long it has been going and how far it has got.
A run that writes to the document takes it read-only until it finishes, and the document says so at the top, with the clock still running: a run ends by re-reading the document, which would discard anything typed underneath it. Only one run goes at a time.
A run outlives the page that started it. Close the tab or reload it and the service carries on. Reopen the document and the run is picked up where it got to, editing still paused, and its result lands as usual. A run that has been finished and collected, or that is more than fifteen minutes past, is gone, and the document simply opens with whatever it wrote. Stop, on the banner or in the dialog, ends a run. It takes effect at the next point the work reports progress, and a run part-way through writing always finishes writing, so a stop never leaves half a document. Whatever it wrote stays.
If the page loses contact with a service, because the connection dropped or the service went quiet for a long time, it says so rather than reporting the run as failed. The service is still working. Reload the document to pick the run back up. When the server does not answer one of a run’s writes, the run says that the change may or may not have been saved.
There are two integration points, both in the Text Editor (see Entering and Tokenizing Text), and Parse is also in the Annotate toolbar:
-
Tokenization: splits the saved text into sentences, tokens and words. The built-in uses the browser’s own segmentation and is always available. A service can replace it with segmentation of its own.
-
Parsing: fills in lemmas, parts of speech, features and dependency relations. There is no built-in, so this needs a connected service. The bundled example wraps Stanza. Where the service offers a language option, it opens on the project’s language, and sentences a person has annotated are left alone unless you say otherwise. A parse writes the document a group of sentences at a time. If it stops partway, each sentence has its previous annotation or its whole new one, and the message says what was written. Parse again to finish: a sentence with no words yet is tokenized and parsed then.
Every annotation and relation a service produces is stamped unverified, in the same violet state as any other machine value (see Machine vs. Human Annotations), so you stay in control of what gets confirmed. Where a service draws the sentence and token boundaries, those are not marked. History names the run that drew them. A service reports what it could not do, and a run that changed nothing says so rather than reporting success.
Writing a service of your own is not hard. The Services chapter of the Plaid manual explains how.
10.10. Scripting
From the project’s Access settings, and from your user profile, you can create named API tokens for scripts and services. A token carries your permissions and does not expire, and actions taken with one are attributed to it in the history.
For writing your own scripts against a project, see the Python client and JavaScript client references.
11. Keyboard Shortcuts
Every shortcut that takes Ctrl also takes Cmd on a Mac. The shortcuts below are fixed and cannot be rebound. The editor carries a short version of this under Gestures and marks, above the first sentence.
11.1. Annotation grid
| Shortcut | Action |
|---|---|
Tab / Shift+Tab |
Move to the next or previous cell in the row. |
↑ / ↓ |
Move between rows. ← / → move along the row from the ends of a value, following the words: in a right-to-left sentence ← moves forward. See Text Direction. |
Enter |
Commit the cell, taking the highlighted suggestion if the list is open. |
Esc |
Cancel the edit, or close an open list or panel. |
Backspace (Features) |
On the empty input, select the last feature, so a second press removes it. |
Alt+↓ |
In a lemma, XPOS or Features cell, list what the project has given words like this one, with counts, and pick one. See What This Project Has Said Before. |
Ctrl/Cmd+Enter |
Accept everything proposed on the current word, then move to the next word. |
Ctrl/Cmd+Backspace |
Discard the machine’s unconfirmed work on the current word, then move to the next word. Anything a person entered or confirmed stays. |
Ctrl/Cmd+Shift+↑ / ↓ |
Jump to the previous or next word that still needs a look. |
Alt+click a word |
Open that sentence in the Text Editor. |
↑ out of the top row |
Jump to this word’s relation label in the tree. |
11.2. Dependency tree
| Shortcut | Action |
|---|---|
Drag from one word to another |
Draw a relation, or re-point an existing one. |
Ctrl/Cmd+drag from one word to another |
Add a relation to the enhanced graph. Over a relation of the tree, give it a different label there. |
Click a label |
Edit it. |
Ctrl/Cmd+click a relation of the tree, or Ctrl/Cmd+E on its label |
Leave it out of the enhanced graph, or put it back. |
Ctrl/Cmd+D |
Jump from the grid into the relation labels. From the tree’s labels, move to the enhanced relations under the words, and from there back. |
← / →, or Tab / Shift+Tab |
Move between labels. Enter edits the one you are on, ↓ drops back into the grid. |
Tab / Shift+Tab in an open label |
Commit it and open the next or previous label. |
Esc |
Leave the labels, or cancel an edit. |
Shift+Backspace in a label you have just opened |
Delete the relation. Once you have typed, it is an ordinary Backspace. |
11.3. Text Editor
| Gesture | Action |
|---|---|
Click a token |
Start or end a sentence there. |
Hover a token |
Open its panel, to edit its words or delete it. |
Alt+click a token |
Open that sentence in the Annotate tab. |
Select a stretch of text |
Make a token of it. |
11.4. Search
| Shortcut | Action |
|---|---|
Enter (word box) |
Write the pattern for what you typed and run it. |
Ctrl/Cmd+Enter (pattern box) |
Run the pattern. |
Tab (pattern box) |
Indent, rather than leaving the box. |