Evaluating a classification model is how you check whether it is labeling conversations correctly. Tekst walks you through a short guided flow: you pick a batch of real conversations, confirm or correct the label the model predicted for each one, and Tekst reports the resulting accuracy.
A batch is ten conversations and takes around five minutes to work through.
Before you start
Make sure you have:
- Access to the Models section of the Tekst platform.
- A classification model that already has its labels configured. The evaluate action stays disabled until then, with a tooltip telling you to add labels first.
- At least one inbox linked to the model, so Tekst has real conversations to sample.
If you have not built a model yet, start with Set up your first classification model.
Start an evaluation
- Go to Models and open the classification model you want to evaluate.
- In the header, find the Model accuracy figure.
- Select Evaluate model.
Tekst opens the Model evaluation panel, which shows the three steps ahead of you and an estimate of how long they take.
Select Start to begin.
Once a model has been evaluated at least once, the header shows the accuracy percentage instead, and selecting it opens the Model accuracy panel, where you can start another batch. If the model has an unpublished draft, the figure shows the draft's score with the published score in brackets, for example 90% (Published: 82%).
A Review N examples pill appears next to the figure when examples are waiting for a correction. See "Examples that need a correction" below.
Step 1: Choose the conversations to review
In the Selection step you decide which conversations Tekst should pull in.
- Inboxes - limit the sample to specific inboxes, or leave it as All inboxes selected.
- Tags - narrow the sample to conversations carrying particular labels. The default is All Tags. Opening the dropdown lists the models linked to the selected inboxes, and expanding one lets you search its labels and pick individual ones. Parent and child labels are shown together, so you can target a specific branch of your label hierarchy.
Select Continue. Tekst samples ten matching conversations and runs the current model on each one, so you have a prediction to check against. If nothing matches your filters, Tekst tells you so and you can adjust them.
Narrowing by label is useful when you want to test one specific part of your label hierarchy rather than the model as a whole. Leaving the filters open gives you a more representative picture of everyday performance.
Step 2: Correct what the model got wrong
You now move through the ten conversations one at a time. Each one is shown with the label the model predicted below it.
- Select Correct if the prediction is right.
- Select Incorrect if it is wrong. The wrong label is struck through in red and a Choose tag picker opens so you can search for the right one. If the label you need does not exist yet, add it with Add new tag.
Once you have picked the right label, the correction is spelled out for you: the wrong label struck through on the left, an arrow, and your chosen label in green on the right. That is your confirmation the correction registered before you move on.
If a conversation was predicted with no label at all, you are asked to choose one directly. You can also Undo a decision if you change your mind.
The numbered squares at the bottom track your progress and let you jump back to an earlier conversation. Your confirmations and corrections become the validated examples Tekst measures accuracy against, so accuracy is only as good as the care you take here.
Skipping and removing
Two options behave differently:
- Skip moves past a conversation without recording anything. Nothing is saved for it.
- Remove from example set, in the menu next to Skip, deletes the conversation from your examples entirely. Use it for a conversation that should never have been an example, not for one you simply do not want to judge now.
When you reach the last conversation, Continue becomes Evaluate model. Select it, and Tekst compares everything you confirmed against what the model predicted. Your progress is saved while this runs, so you can close the panel and come back.
Closing the panel before you finish discards the corrections for the current batch, and Tekst warns you before it does.
Step 3: Read the accuracy results
When the batch finishes, Tekst opens the Model accuracy panel.
At the top is Model accuracy. On a model with an unpublished draft this is two cards, Published model and Draft model, and selecting one switches the breakdown below between them. On a model with no draft it is a single card.
Below it is Tag evaluation, the per-label breakdown. Each label is marked either All correct or with the number of conversations it got wrong. Labels that had nothing to review in this batch are listed without a marker. When your labels are nested, they are grouped under their parent, so you can collapse a branch you are not looking at. In the draft view, a count that moved is struck through with the new number beside it, so you can see exactly what your changes did to each label.
Select any label to open its detail view, which puts the label you confirmed beside the one the model predicted, conversation by conversation. Conversations the model got wrong are listed first and the correct ones are collapsed behind a Show correct toggle, so what needs your attention is what you see. That comparison is the fastest way to spot a label description that is too vague or overlaps with a neighboring label.
Two things to keep in mind when reading the number:
- It reflects the examples reviewed so far, not your whole message history. Reviewing more conversations can move it up or down without the model itself having changed.
- Select Add more examples at the bottom of the panel to sample another batch and build a firmer picture.
For the detail of how the score itself is computed, see How classification accuracy is measured.
Let Tekst improve the descriptions
When the results show room to improve, select Improve accuracy at the bottom of the Model accuracy panel. Tekst uses the examples you have confirmed or corrected to work out better label descriptions.
The run works in the background and can take up to 30 minutes, so you can close the panel and pick the result up later. Anything it finds lands in a draft, which means you read the new descriptions and decide whether to publish them. Nothing reaches live messages until you do.
For the full walkthrough, including what each outcome means and how to cancel a run, see the "Improve model accuracy automatically" article.
Where your corrections are kept
Everything you confirm during an evaluation is stored as an example on the model. To see or edit those examples later, open the Model accuracy panel from the model header and switch to the Examples tab. There you can review each conversation, change the confirmed label, filter the list by label, and remove examples you no longer want counted.
Add an example from the Messages page
You do not have to run an evaluation to grow your example set. As messages arrive you can review them on the Messages page and add the good ones as you go.
- Go to Messages and open a conversation.
- In the side panel, open the Tags tab. It lists the labels the model assigned, grouped by the model that assigned them.
- Check the label. If it is right, mark it Verified. If it is wrong, change it to the correct one first.
- Select the bookmark icon, Add as example.
The conversation joins the same example set the evaluation flow builds, so it counts toward the model's accuracy and feeds future improvements. This is a good habit while a model is new: verify a handful of messages as you work through them and your example set grows without a separate review session.
Examples that need a correction
If you change a model's labels after building up examples, some of those examples no longer line up with the new configuration. Tekst marks these as needing correction and leaves them out of the accuracy figure until you revisit them. A Review N examples pill appears next to the accuracy figure in the model header, and the Model accuracy panel prompts you with Add missing corrections.
These corrections also gate publishing. Selecting Publish while examples are waiting opens a panel instead of publishing, so the change cannot go live against an example set the model can no longer answer.
From there, Review examples walks you through the corrections, and Remove examples and publish drops those examples and publishes straight away. See the "Publish a model draft" article for the full picture.
Related articles
- What is a classification model?
- Set up your first classification model
- How classification accuracy is measured
- What data does Tekst store and how does the learning process work?
- To understand the draft workflow behind Publish, see the "What is a model draft?" article.
- For automatic description rewrites, see the "Improve model accuracy automatically" article.
0 comments
Please sign in to leave a comment.