Shift Left, Test Right seriesShift Left, Test Right · Chapter 3 of 3
Jira Guides17 min read

How an MCP Took Our Real Coverage From 65% to 87%

We had hundreds of requirements and test cases living next to the code, where nothing could report on them. An agent moved them into the catalogue, checked our homework, and found more than we expected.

Balázs Szakál

Balázs Szakál

Founder & QA Lead at BesTest

Updated September 23, 2026
How an MCP Took Our Real Coverage From 65% to 87%

Disclaimer: as in the first two chapters, I used AI to help formatting my thoughts™ - the opinions, the numbers and the Hungarian sentence structure are all mine. The numbers in particular are all mine, which is rather the point of this chapter. They are from this summer; I held the post back until the feature it is about was ready for you to try.

At the end of Chapter 2 I promised we would come back to requirements. This is that chapter, and it took a detour through our own product.

We are testers. We had written hundreds of requirements and test cases for BesTest over the years, and most of them were not in BesTest. They lived next to the code: feature documents in the repository, BDD scenarios, end-to-end test files, and a fair amount in people's heads. That is a perfectly good place for a developer to read them. It is a useless place to report on them. You cannot open a dashboard on a Git folder and ask which feature is under-tested.

We tried to move them into the catalogue the normal way, with exports and imports and typing. It was an enormous amount of work to keep it updated, so it stalled, and for four months our own project barely moved.

Then our MCP server arrived, and we pointed an agent at the repository and said: read this, create the requirements, link the tests. Three weeks later we had a catalogue we could measure. The first measurement said 87% covered. When we asked the harder question - is there a test on the other end that actually runs and proves this - the answer was 65%.

A month of sweeping later, the features we had been through sat at 87% to 100%, with a running test behind every point. This chapter is about what the agent found in between, and why the problem it solved is hard for every team, not just a careless one.

First, what an MCP server actually is

Skip this section if you already know. I am including it because when I explain this to QA people, most have heard the acronym and almost nobody has tried it. The explanations online are written for developers, and in a lot of companies the AI policy has not yet caught up to the point where a tester is sure they are even allowed to plug an agent into anything. So here is the whole idea, in plain language.

You have an AI assistant. On its own, it knows nothing about your company. It cannot see your Jira, your requirements, or your test results. If you want it to help, you paste things into the chat window and hope you pasted the right things.

MCP (Model Context Protocol) is a standard way to give an assistant a door into a real system. Somebody runs a small server that speaks MCP and knows how to talk to one product. You point your assistant at that server. Now the assistant can look things up itself, and change things, without you pasting anything.

That is genuinely it. It is a plug shape. Atlassian publishes one for Jira. We publish one for BesTest. Dozens of other products publish their own.

The reason this matters for testing specifically: you can now ask a question in plain language that would previously have been a reporting project.

Not "export the traceability matrix to CSV, open it in Excel, pivot it, and eyeball the empty cells", but:

> "Which requirements in this project have no test that actually executes?"

And the assistant goes and finds out. It reads the requirements, follows the links, checks what is on the other end of each one, and comes back with a list. In our case it came back with a list we did not enjoy reading.

One more thing worth knowing, because it is the part people get wrong: an MCP server is not an AI. It has no opinions and does not generate anything. It is a doorway. All the thinking happens in the assistant, which can be whichever one your company allows: a cloud model, a local one, a terminal tool, a chat window. The server just makes your data reachable, under your own permissions.

Why one server was not enough

Atlassian has its own Jira MCP server, and it is good. Connect it and your agent can read your sprint, your work items, your comments.

But it cannot see test coverage. Not because Atlassian did something wrong - because in our setup the test data is not in Jira. As Chapter 2 explained at length, we deliberately do not store test cases as Jira work items. So when an agent asks Jira "what is tested here", Jira has nothing to say. It does not know what a test case is.

So we ran both. Atlassian's server for what is in the sprint, ours for what is covered, and the agent joins them on the work item both sides already reference.

That combination is what makes the interesting question askable:

> "Which stories in this sprint have no test coverage?"

Neither server can answer that alone. Jira knows the sprint and not the coverage. We know the coverage and not the sprint. Together it is one question and about four seconds.

Five things you can ask an agent once Jira and BesTest are both connected over MCP
Five things you can ask an agent once Jira and BesTest are both connected over MCP (click to enlarge)

These are the examples from our AI testing pages, where each one is written up with what the agent does behind the prompt. The sprint one is the combination above.

If you want the practical setup, we wrote it up separately in Jira MCP server setup, including Atlassian's endpoint and what it can and cannot see. Trust Atlassian's docs over ours for their half.

Why our own catalogue was frozen for four months

The sequence matters, because the tidy version would leave out the part that is useful.

We did not begin with a plan to audit our coverage. We began because our own test management was thin, and we knew exactly why.

We build a test management tool. Our requirements and tests existed, but they lived in the repository: feature documents next to the code, BDD scenarios in feature files, end-to-end tests with keys in their names. Developers could read them. Nobody could report on them. When I wanted to know whether a feature was properly covered, I had to open something that manages Git and read, which is not what I want to do to answer a management question on a Tuesday afternoon.

We wanted all of it in the catalogue, where coverage is a number and a gadget rather than a reading exercise. We tried. Exporting, reshaping, importing, hand-creating: it was an enormous amount of time for something that was never the most urgent thing. By the end of June our own project contained:

  • •49 requirements
  • •76 test cases
  • •70 links between them
  • •58 test executions, in a single test cycle, ever

That is not a test suite. That is a demo dataset with a job title.

And the detail that still bothers me: those numbers did not move between the start of March and the end of June. Four months. Exactly 70 links, exactly 49 requirements, unchanged, while we shipped feature after feature. The knowledge was there the whole time. It was just in a place that cannot count.

The MCP is what broke that, and not for a noble reason. It broke it because typing became unnecessary. Instead of opening the app and creating 40 requirements by hand, we could point an agent at a feature's documentation and say "read this and create the requirements". It read the spec, created them, linked them to the tests it could identify, and told us which ones it could not.

An agent reading a feature document and creating the requirements in BesTest over MCP
An agent reading a feature document and creating the requirements in BesTest over MCP (click to enlarge)

Three weeks later the same project held:

End of JuneEnd of July
Requirements49535
Test cases76800
Requirement-to-test links701,672
Test executions583,823
Test cycles19

I am not going to pretend that is a productivity miracle. Most of it is transcription: information that already existed in specs and in people's heads, moved into a place where it could be counted. The miracle, if there is one, is that transcription stopped being expensive enough to keep postponing.

And then we could measure. Which is where it got interesting.

The number that looked fine

With a real catalogue in place, we ran the obvious query: how many of our requirements are covered?

87.1%. 466 of 535 requirements had at least one test case linked.

Our own dashboard in July: 466 of 535 requirements covered, 87.1%
Our own dashboard in July: 466 of 535 requirements covered, 87.1% (click to enlarge)

Twelve days earlier the same query had said 79.2%. So the number was going up and to the right, we had a graph, and I very nearly wrote this chapter about that graph.

The reason I did not is that "covered" was doing an enormous amount of work in that sentence.

Ask yourself what your own coverage number actually asserts. In most tools, including ours by default, it asserts this: a link exists between this requirement and at least one test case. That is all. It does not assert that the test on the other end runs. It does not assert that the test passes. It does not assert that the test has anything to do with the requirement it is attached to.

It is a measure of whether somebody did the paperwork.

So we asked a harder question, and this is exactly the sort of question that used to be too expensive to ask:

> "For every requirement, do not tell me whether a link exists. Tell me whether there is a test on the other end that actually executes in our suite."

That is a different query. It has to follow every link, look at what is on the far end, and work out whether that thing is a real automated test or merely a document describing one.

The answer was 299 of 461 real requirements. 64.9%.

Two coverage numbers, 87% reported and 65% proven, both correct
Two coverage numbers, 87% reported and 65% proven, both correct (click to enlarge)

Hollow: 89 requirements with a story and no proof

The gap between 87% and 65% has a shape, and once we had a name for it we could not stop seeing it.

We ended up with three categories:

  • •Covered - the requirement is linked to at least one test that genuinely runs. 299 requirements.
  • •Hollow - the requirement has links, and every single one goes to a BDD scenario. Not one goes to an executable test. 89 requirements.
  • •Gray - no links at all. 73 requirements.

Gray is the easy one. It shows up as a gap on every dashboard, it comes up in a review, and somebody eventually fixes it. Gray is a problem that behaves like a problem.

Hollow is the dangerous one. A hollow requirement renders as covered. It is green. It has a link, and if you click the link there is a real document at the end of it, written in Given/When/Then, describing precisely how this requirement should behave.

Nothing runs it.

The phrase I wrote in our internal notes and have not been able to improve on is: a scenario describes it, nothing proves it.

This is a specifically modern failure. BDD scenarios are good. Writing behaviour down in structured language before you build it is good. But a scenario is a description, and at some point somebody has to connect the description to an actual executing test, and if that second step quietly does not happen, you are left with a document that looks exactly like coverage from every angle except the one that matters.

We had 89 of those. On our own product. While selling coverage reporting. I do not think that makes us unusual; I think it makes us a team with BDD scenarios.

A requirement with significance High, five to seven test cases expected, and zero linked
A requirement with significance High, five to seven test cases expected, and zero linked (click to enlarge)

This is the mechanism, on a demo project rather than our own: significance is derived from complexity times impact, and it sets how much coverage the requirement needs before it counts as covered. The requirement above expects five to seven test cases and has none, so it reports a gap. A hollow requirement is the harder case, because the "actual" count is not zero. It is one, and the one is a scenario.

Before anybody concludes that BDD was the villain: we tested that theory and it was wrong. We re-ran our own risk scoring with every BDD row stripped out of the calculation, expecting the risk picture to change substantially. It moved from 188 high-risk items to 186. Two. Deleting more than half the corpus would have changed essentially nothing about which parts of the product are actually risky.

So we kept BDD. What we changed was how a requirement gets to be green. A requirement can now be marked Covered by hand. That toggle overrides the calculated coverage everywhere it is shown - the catalogue, the gadgets, the reports - so a requirement that is proven somewhere the catalogue cannot see (a unit test, a manual check, a scenario that genuinely is enough on its own) turns green because a person decided it should, with their name on the decision. Green by accident is gone. Green is either a running test or a deliberate call.

The Covered toggle on a requirement: a manual override that counts as covered in widgets and the significance matrix regardless of link count
The Covered toggle on a requirement: a manual override that counts as covered in widgets and the significance matrix regardless of link count (click to enlarge)

The requirement that was simply false

This is the finding I did not see coming, and it is the reason I wanted to write this chapter at all.

One of our requirements, in the test-cases feature, stated the validation rules for creating a test case. Name between 3 and 255 characters. Description up to 5000. Objective up to 2000. Preconditions up to 2000. Specific, plausible, the kind of requirement nobody re-reads.

It was false. Not out of date. False, and as far as we can tell it was never true.

What the product actually enforces when you create a test case is: the name is not empty, and any required custom fields are filled in. There are no length checks at all. The specific numbers in the requirement turned out to be a garbled copy of the rules from our import validators, which are a different code path with different limits.

So how does a false requirement survive in a project that has tests?

Because it had a test. A diabolical 666-line test file that asserted, in loving detail, that the validation rules worked exactly as the requirement described. The test passed every single time.

The test passed because it did not test the product. It contained its own inline copy of the validator logic and then asserted against that copy. It was checking that a function it had just defined behaved the way it had just defined it. It never imported anything from the application. The same file also tested a planned-time parser that accepts input like "1h 30m". We have never shipped that parser. It does not exist.

This is what generated test code looks like when nobody reads it. It is fluent, it is long, it is green, and it proves nothing. Agents will produce this at volume if you let them, and the volume is exactly the point of using them, so the fix cannot be "generate less". The fix is that nothing generated gets to count until a person has looked at it. That is why review is built into BesTest: a test case an agent creates lands as a draft, goes to a named reviewer, and only an approved test case is part of the project. The same review flow covers requirements. An agent can write the requirement, an agent can write the test, and a human still decides whether either of them is true.

FXGT-TC-419, created by an agent, waiting in review with a named reviewer and the note "Agent created, user acceptance needed"
FXGT-TC-419, created by an agent, waiting in review with a named reviewer and the note "Agent created, user acceptance needed" (click to enlarge)

Somebody had already flagged that this test was tautological. What nobody did was check the other direction. Everyone assumed the worst case was "this test proves nothing". The actual worst case was worse:

A tautological test does not merely fail to test something. It can manufacture a requirement.

The test asserted rules. The rules got written down as a requirement. The requirement pointed at the test as its evidence. The loop closed, and for months our documentation described a product that did not exist, with a green checkmark next to it.

If you take one thing from this chapter, take this: when you find a test that only tests itself, do not just delete the test. Go and re-read whatever that test was claiming to cover, because it may have invented it.

Live preview - this is the real BesTest UI

The number on your dashboard is answering an easier question than you think.

Most coverage reports, ours included by default, count a requirement as covered when something is linked to it. That is a measure of paperwork, not of proof. BesTest keeps requirements, test cases and executions in one place inside Jira so the harder question is actually askable, and the MCP server that made this audit possible is in beta: create your own token from the app menu and connect any agent. Free for up to 10 users, about a minute to switch on.

View on the Atlassian Marketplace

Tests nobody claimed, and the mess retirement leaves

Two smaller findings, both of which generalise.

Orphans: real tests that no requirement names. We found 24 automated tests that run, pass, and prove something useful, and that no requirement points to. Their coverage is real and completely invisible to every report we have. That number is down from 62 when we started looking, which means most of it was fixable paperwork, but 24 is still 24.

My favourite two are called BT-E2E-ONB-10 and BT-E2E-ONB-11. They were built during an earlier onboarding improvement, specifically to close a coverage gap. They work. Nobody added them to a matrix, so the very gap they were built to close still read as open. The work was done and the scoreboard never noticed.

Retirement leaves residue, and it hollows out other people's work. We retired an old peer-review feature and replaced it with a rebuilt one. We archived its 18 test cases and its 11 requirements. Clean.

Except two requirements in a different feature had been linked to two of those archived tests. Those requirements still read as covered. Nothing behind them could ever execute again, because the tests they pointed to had been removed from the registry entirely. The retirement quietly punched a hole in a neighbouring feature and nothing anywhere reported it.

We swept for this pattern and found more of it, including 11 links running from a single Jira issue to requirements that no longer exist, and 53 link rows in total attached to something archived on one end or the other.

None of that is exotic. Any team that has ever deprecated a feature has some version of this. The reason nobody finds it is that finding it requires walking every link and checking whether the far end still exists, which is precisely the kind of tedious full-corpus question that is trivial to ask an agent and miserable to do by hand.

From 65% to 87%: what the sweep actually did

We went feature by feature, 26 of them. Each one got the same treatment: every requirement gets a single explicit verdict, not just a link.

The verdicts we allow are deliberately blunt:

  • •automated - name the test keys, and note any caveat
  • •covered elsewhere - a unit, service or database test covers it; name the file
  • •manual only - genuinely needs human eyes or a real Jira instance
  • •not e2e testable - architecture or non-functional; say so
  • •gap - should be automated, is not
  • •verified once - proven at the release that shipped it, never re-run since

And then coverage is one formula with no wiggle room in it: automated / (automated + gap). Everything else has to be argued for in writing, or marked Covered by a person who takes responsibility for it.

The first three features through the sweep:

FeatureCoverage, with proofMatrix rows naming a real test
test-collections90% (9 automated, 1 gap, 1 not e2e testable)1 → 9
test-cases87% (13 automated, 2 gaps)1 → 13
issue-panel100% (13 automated, 1 covered elsewhere)11 → 13

That is the 87% in the title. It is the same number the dashboard had shown in July, except now there is a running test behind every point of it, and the two gaps are named rather than hidden inside a green bar.

Now the part that surprised me most, and that I would not have believed if somebody had told me in June.

We closed 12 gaps across those features. Number of new end-to-end tests written: zero.

Every single gap was closed by adding a step to a test that already existed, or by one small database-level test. The reason is a rule we adopted early in the sweep and have not regretted once: before you write a new test, you must search the entire suite and explain in writing why an existing test cannot be extended. A new test has to earn itself. It needs a genuinely different setup, or an assertion that can only be made once, or a failure that has to be attributable on its own.

Nearly every time, the answer was that a test already arranged exactly the right state and simply was not asserting the thing we cared about.

I will not oversell the result. All 12 additions passed on the first run, which means not one of them caught a live defect. What changed is that a dozen guarantees which had been holding by accident are now defended on purpose - including two whose failure would have been silent and expensive.

There is a real lesson in the suite not growing. Our coverage problem was never a shortage of tests. It was a shortage of knowing what the tests we already had were for.

What the beta looks like from our side

Our own project is one data point, and I am aware it is a flattering one. So before publishing I looked at the whole fleet, every team on the beta across all regions, to see whether the pattern held for anybody who is not us.

It does, and the shape of it is more interesting than the size.

Since the beta opened in mid-July, roughly a third of the active teams have created an API and MCP token. That third is responsible for about three quarters of all the requirements written in the product in that period, and for more than half of the links between requirements and test cases. Teams without a token write test cases too, sometimes a lot of them at once. What they mostly do not write is requirements, and what they mostly do not do is link.

Share of everything created on the beta by the third of teams with an agent connected, and the same share for our own project
Share of everything created on the beta by the third of teams with an agent connected, and the same share for our own project (click to enlarge)

I do not want to overclaim from a beta with a handful of teams on it, so read this as a shape, not a statistic. But the shape matches exactly what happened to us: the expensive half of test management was never the test case. It was the requirement, and the wire from the requirement to the test. That is the half nobody has time for on a Tuesday, and it is the half that stops being expensive when an agent does the typing.

For our own project the share is over ninety percent for every entity type. Nearly everything in our catalogue since July arrived through the same door this chapter is about.

There is a bigger point behind that number. Before AI, writing test cases and requirements at this volume was expensive enough that a small team could not afford a real catalogue; you wrote what you had time for and the rest stayed in your head. That constraint is gone. If you have the need, you can now produce hundreds of items in a week, and once you can, you are not going to maintain them by clicking through forms. You will manage them programmatically, through a REST API or through an agent over MCP. That is the reason we built both.

What else the audit turned up

The three features above were the first ones through. The rest of the sweep found more, and it is worth listing, because every one of these patterns will exist somewhere in your project too.

custom-fields was our weakest area, and it had been marked as done. 56 requirements, 30 of them hollow - a third of every hollow requirement in the entire project, concentrated in one feature an earlier audit had already signed off. Only 6 of its 56 traceability rows named a test that actually runs; most of the rest named unit test filenames, which is a different kind of coverage wearing the same badge. It also carried the largest population of BDD scenarios anywhere in the corpus, which is precisely how it got that way. It went back through the sweep with the full verdict-per-requirement treatment.

dashboard-gadgets owned 37 automated tests and could not tell you what they proved. Second-largest test population in the product, and 5 of its 12 requirements had no links at all. There were no orphans, so those 37 tests were claimed by something; the requirements just did not name them. Plenty of testing, almost no traceability. The fix was paperwork, and an agent did most of it.

import-custom-fields had 6 traceability rows for 16 live requirements. Ten requirements appeared nowhere in its matrix. Same treatment.

Three features we had already marked as audited had not really landed. They were done under an earlier, weaker version of the pass that only checked whether a link existed rather than demanding a verdict per requirement. They passed that bar and were in worse shape than several features we had not audited at all. We redid them.

One clarification so this is not misread as worse than it was: our project also contains 26 decision-log entries which show as having zero coverage. That is deliberate. They are records of decisions, not testable behaviour, and they are excluded from every coverage figure in this article.

For the record, exactly one feature was complete before the sweep touched it: requirements. 14 of 14, no gaps, no hollow rows, no orphans. One out of twenty-six.

The traceability matrix for the requirements feature: 14 of 14, every test green
The traceability matrix for the requirements feature: 14 of 14, every test green (click to enlarge)

We were not amateurs before this. We were a team without a way to see the whole picture at once, which is the normal condition, and which is what the MCP changed.

The five findings: 89 hollow requirements, one false requirement, 66 times more executions, four frozen months, and zero new tests written to close twelve gaps
The five findings: 89 hollow requirements, one false requirement, 66 times more executions, four frozen months, and zero new tests written to close twelve gaps (click to enlarge)

How to run this on your own project

You do not need our product to do most of this. You need a way to ask questions across your whole corpus without a week of spreadsheet work, and you need to be willing to accept an unflattering answer.

The four questions that produced everything above:

  • •"For each requirement, is there a test on the other end of its links that actually executes?" This is the one that finds hollow. If your tooling can only tell you whether a link exists, you do not currently know your coverage.
  • •"Which tests run and pass, but no requirement names?" This finds orphans, and it usually finds work your team already did and never got credit for.
  • •"Which links point at something that has been archived or deleted?" This finds the residue of every retirement you have ever done.
  • •"Which tests never import anything from the application?" This finds tautological tests. When you find one, re-read the requirement it claims to cover, because it may have written it.

Question 4 is the one I would run first if I were starting over. It is the cheapest to check and it was by far the most alarming thing we found.

If you want to try it inside Jira with our tooling: the BesTest MCP server and REST API are in beta, and you set them up yourself. Open the app menu, choose API & MCP tokens, create a token, pick the Space it reaches and whether it is read-only or read-and-write, and paste it into whichever agent you use. No request form, no waiting on support. A Space admin can see and revoke every token pointed at their Space. The setup docs have the exact steps for the common agents.

The API & MCP tokens screen: your own tokens, one Space each, read-only or read-and-write, with expiry and last use
The API & MCP tokens screen: your own tokens, one Space each, read-only or read-and-write, with expiry and last use (click to enlarge)

The MCP is what made all of the above askable in an afternoon rather than a quarter, and I would rather you used it and found something embarrassing than looked at a green dashboard for another four months.

What this has to do with requirements

I promised a chapter about requirements, and this is the part where it turns out it was one all along.

Writing requirements is not the hard part. Keeping them alive is: a list that still matches the product, where every item is either proven by a test that runs or deliberately marked as covered by someone who checked, and where a test case that proves nothing does not get to hide behind a link. Both halves matter. A requirement with no test is a gap. A test with no requirement is effort nobody can see. Our reports, gadgets and catalogues exist to make both of those visible, to your eyes or to your agent.

What changed this summer is who maintains the list. Our requirements catalogue is now living documentation that gets updated by talking to it. Read this feature document and create the requirements. Which of these are stale after the change in PAY-42? Which test cases have found the most defects, and which ones fail most often? Draft ten test cases for the uncovered paths and send them to review. The agent does the walking; BesTest stores the result and shows the report; a person approves what gets to count.

You can do this with any agent you already pay for, or a local model, because the MCP is just a door. You manage your own tokens, per person, per Space, read-only or read-and-write. Your test cases can be traditional steps, BDD, imported, or end-to-end; linking works the way linking works in Jira; and everything created by an agent goes through the same review as everything created by a human, with a named reviewer who does not need a terminal to do their job. Your team talks to the agent, the agent talks to BesTest, and the review is where the two meet.

And if you run question 4 on your own suite and it finds something, I would genuinely like to hear about it.

Frequently Asked Questions

What is an MCP server, in plain language?

MCP (Model Context Protocol) is a standard way to give an AI assistant a door into a real system. Someone runs a small server that speaks MCP and knows how to talk to one product, you point your assistant at it, and the assistant can then look things up and make changes itself instead of you pasting data into a chat window. An MCP server is not itself an AI: it has no opinions and generates nothing. All the reasoning happens in the assistant, which can be any cloud or local model your company allows. The server just makes your data reachable, under your own permissions.

Why did your coverage show 87% when the proven figure was 65%?

Because the two numbers answer different questions. The 87% figure counts a requirement as covered when at least one test case is linked to it, which is what most coverage reports measure. The 65% figure only counts a requirement as covered when there is a test on the other end of the link that actually executes. The 22-point gap was 89 requirements linked exclusively to BDD scenarios: documents describing the expected behaviour, with nothing running them. They render as covered on any link-based report. After the sweep, the features we had been through sat at 87% to 100% with a running test behind every point.

What is a hollow requirement?

A requirement that has at least one link, but where every link goes to something that describes the behaviour rather than something that tests it, typically a BDD scenario with no automated test wired to it. It is more dangerous than a requirement with no links at all, because a requirement with no links shows up as a gap and eventually gets fixed, whereas a hollow requirement renders as covered and green on every dashboard. In BesTest a requirement can be marked Covered by hand, which overrides the calculated coverage in every report, so green is either a running test or a deliberate, named decision.

How can a test create a false requirement?

If a test contains its own inline copy of the logic it claims to verify, it passes without ever touching the application. We had a 666-line test file that asserted validation rules the product does not implement, including a parser we have never shipped. Those invented rules were then written down as a requirement, and the requirement cited the passing test as its evidence. The loop closed. When you find a test that only tests itself, re-read whatever it claimed to cover, because it may have manufactured it. This is also why BesTest routes agent-created test cases and requirements through review before they count.

Did closing your coverage gaps require writing many new tests?

No. We closed 12 gaps across three features and wrote zero new end-to-end tests. Every gap was closed either by adding an assertion to a test that already existed or by one small database-level test. We apply a rule that a new test must be justified in writing against the whole existing suite first, and nearly every time an existing test already arranged the correct state and simply was not asserting the thing we cared about. The shortage was never tests. It was knowing what the existing tests were for.

Can an AI agent see test coverage through the Jira MCP server?

Not if your test data lives outside Jira work items. Atlassian's Jira MCP server can tell an agent what is in a sprint, who is assigned, and what the comments say, but it has no concept of a test case unless test cases are stored as Jira issues. Running a test management MCP server alongside it lets an agent join the two on the work item both sides already reference, which is what makes questions like "which stories in this sprint have no test coverage" answerable.

Do I need to ask support for a BesTest MCP key?

No. The BesTest MCP and REST API are in beta and fully self-service. Open the app menu in BesTest, choose API & MCP tokens, and create a token yourself: give it a name, pick the Space it may reach, choose read-only or read-and-write, and set an expiry. The full value is shown once, so paste it into your agent's configuration straight away. You can rotate or revoke it at any time, and a Space admin can see and revoke every token pointed at their Space.

Tags:MCPAI agentstest coveragerequirements traceabilitydogfoodingtesting in jirashift left

The number on your dashboard is answering an easier question than you think.

Most coverage reports, ours included by default, count a requirement as covered when something is linked to it. That is a measure of paperwork, not of proof. BesTest keeps requirements, test cases and executions in one place inside Jira so the harder question is actually askable, and the MCP server that made this audit possible is in beta: create your own token from the app menu and connect any agent. Free for up to 10 users, about a minute to switch on.

View on the Atlassian Marketplace
Balázs Szakál

Balázs Szakál

Founder & QA Lead at BesTest

Founder of BesTest and QA professional with extensive experience in software testing, test management, and Jira administration. Built BesTest to give testing teams complete visibility from requirements to release.

More about the team →
Getting started

Live in about a minute.

  1. ~30 seconds
    1.Install from the Marketplace

    One click on "Get it now" - no sales call, no signup form, no separate login.

  2. ~1 minute
    2.Enable it on a Space

    Flip it on in Space settings. BesTest shows up in the Space menu, where your team already works.

  3. right away
    3.Run your first test

    Create a requirement, link a test case, hit run. No training course required.

Host your data in the EU, US, or IndiaNo Jira issue bloat - your library stays out of Jira’s wayBuilt on Atlassian Forge
In development · No Jira required

We are taking BesTest out of Jira

Most of what makes BesTest good was never really about Jira. The requirements, the test cases and the coverage model are ours. A standalone version is in the works for teams who do not run Jira at all.

  • The same requirements, test cases and coverage engine
  • Cloud-hosted, in the EU, US or India
  • One email at launch, nothing else
What we can and cannot promise yet

Want to know when it opens?

By joining you agree that Beard & Tailor Studio Kft. may store your email address to send you one email when the standalone version launches. We do not share it with anyone. Ask us to delete it at any time at info@btstudio.io. See our Privacy Policy.