Evaluating an extraction model is how you check whether it is pulling the right values out of your conversations and documents. Tekst walks you through a short guided flow: you pick a batch of real conversations, confirm or correct the values the model extracted from each one, and Tekst reports the resulting accuracy and offers improvements you can apply.
A batch is five conversations and takes around five minutes to work through.
Before you start
Make sure you have:
- Access to the Models section of the Tekst platform.
- An extraction model that already has its entities configured. The evaluate action stays disabled until then, with a tooltip telling you to add entities first.
- At least one inbox linked to the model, so Tekst has real conversations to sample.
If you have not built a model yet, start with Set up your first extraction model.
Start an evaluation
- Go to Models and open the extraction model you want to evaluate.
- In the header, find the Model accuracy figure.
- Select Evaluate model.
Tekst opens the Model evaluation panel, which shows the three steps ahead of you and an estimate of how long they take.
Select Start to begin.
Once a model has been evaluated at least once, the header shows the accuracy percentage instead, and you can reopen this flow from there. A pill next to the figure tells you when something needs your attention, for example Review evaluation when results are waiting or Evaluating while a run is in progress.
Step 1: Choose the conversations to review
In the Selection step you decide which conversations Tekst should pull in.
- Inboxes - limit the sample to specific inboxes, or leave it as All inboxes selected.
- Tags - narrow the sample to conversations carrying particular labels. The default is All Tags. Opening the dropdown lists the models linked to the selected inboxes, and expanding one lets you search its labels and pick individual ones.
Select Continue. Tekst samples five matching conversations and runs the current model on each one, so you have extracted values to check against. If nothing matches your filters, Tekst tells you so and you can adjust them.
Filtering by label is a practical way to focus on one document type. If a classification model already labels incoming mail by document type, you can use that label to review only the purchase orders, or only the invoices, rather than a mixed sample.
Step 2: Correct what the model got wrong
You now move through the five conversations one at a time. For each one you see the full set of fields the model extracted, with the value it found for each.
- Confirm each field the model got right. To accept the whole result at once, use Mark all as correct. Once you have already confirmed or corrected some of the fields, that button reads Mark remaining as correct instead.
- Edit any value the model got wrong. That includes clearing a value it should not have found at all, and filling in one it missed. A field the model found nothing for shows as Empty.
The icon beside each field name indicates the kind of value that field expects, such as a number or free text. That matters when you read the accuracy later, because numbers are compared exactly while text is compared more loosely.
Work through every field before continuing. Tekst asks you to either correct a field or mark the remaining fields as correct, because a field you never looked at cannot count as validated.
The numbered squares at the bottom track your progress and let you jump back to an earlier conversation. Your confirmations and corrections become the validated examples Tekst measures accuracy against, so accuracy is only as good as the care you take here.
Skipping and removing
Two options behave differently:
- Skip moves past a conversation without recording anything. Nothing is saved for it.
- Remove from example set, in the menu next to Skip, deletes the conversation from your examples entirely. Use it for a conversation that should never have been an example, not for one you simply do not want to judge now.
When you reach the last conversation, Continue becomes Evaluate model. Select it, and Tekst compares everything you confirmed against what the model extracted. Your progress is saved while this runs, so you can close the panel and come back.
Closing the panel before you finish discards the corrections for the current batch, and Tekst warns you before it does.
Step 3: Read the accuracy results
The Evaluation step opens with Model accuracy: a Current model card showing how much of what you reviewed the model got right.
Below it is Entity evaluation, the per-field breakdown. Each field is marked either All correct or with the number of values it got wrong. Fields that had nothing to review in this batch are listed without a marker. When your entities are nested, they are grouped under their parent with a count of the fields they hold, so you can collapse a branch you are not looking at.
Because the score averages across every field, a model can look healthy overall while one field is consistently wrong. The breakdown is what surfaces that, so read it rather than the headline number.
Select a field to open its detail view, expanding its parent group first if your entities are nested. That view puts the value you confirmed beside the one the model extracted, with the field's description and expected format alongside. Values the model got wrong are listed first and the correct ones are collapsed behind a Show correct toggle, so what needs your attention is what you see. It is the fastest way to spot a field that needs a clearer description or a different output format.
Two things to keep in mind when reading the number:
- It reflects the examples reviewed so far, not your whole message history. Reviewing more conversations can move it up or down without the model itself having changed.
- Select Review 5 more conversations to sample another batch and build a firmer picture.
For the detail of how the score itself is computed, including how text values are compared, see How extraction accuracy is measured.
Apply suggested improvements
When the results show room to improve, select Suggest improvements. Tekst uses the corrections you just made to work out better entity descriptions.
This runs in the background and your progress is saved, so you can close the panel and pick the result up later from the model header.
When it finishes you get two cards side by side: Current model and Suggested model, each with its own accuracy. The breakdown below reports the effect field by field, as the count of wrong values before and after the change, so a field reading 5 → 0 incorrect is one the suggestion fixes outright.
Read that list before you apply, because a higher overall figure can hide a field that got worse. A suggestion can lift a model from 70% to 88% while taking one field from none wrong to several. If that field is the one an automation depends on, the trade may not be worth making.
- Apply improvements updates the model. Tekst confirms with Improvement applied.
- Discard improvements throws the suggestion away. This cannot be undone, though you can run a new optimization later.
The button is not offered when the model is already at 100% on your examples, since there is nothing to improve.
If Tekst cannot find anything better, it tells you the extraction is already well optimized and invites you to correct more conversations. That is the right next move: more corrections give the optimization more to learn from.
Where your corrections are kept
Everything you confirm during an evaluation is stored as an example on the model. To see or edit those examples later, open the Model accuracy panel from the model header and switch to the Examples tab. There you can review each conversation, change the confirmed values, and remove examples you no longer want counted.
Add an example from the Messages page
You do not have to run an evaluation to grow your example set. As messages arrive you can review them on the Messages page and add the good ones as you go.
- Go to Messages and open a conversation.
- In the side panel, open the Entities tab. It lists every field the model extracted with the value it found, nested the way your entities are configured, and shows Empty where it found nothing.
- Work down the list, verifying each field and correcting any value that is wrong.
- Select Add as example.
Every field has to be verified before Add as example becomes available. That is deliberate: an example is only useful as ground truth if someone has actually checked all of it.
The conversation joins the same example set the evaluation flow builds, so it counts toward the model's accuracy and feeds future improvements. This is a good habit while a model is new: verify a handful of messages as you work through them and your example set grows without a separate review session.
If you change a model's entities after building up examples, some of those examples no longer line up with the new configuration. Tekst marks these as needing correction and leaves them out of the accuracy figure until you revisit them, which it prompts you to do with Add missing corrections.
0 comments
Please sign in to leave a comment.