KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesReporting a Statistic You Did Not ComputeProject Delivery · Research ProjectsLesson 261/267← PrevNext →
GuidePublished 16 Aug 202615 min readBy KEVOS Editorialreporting statistics in a thesiscorrelation coefficient reportinginterpretation keysafety index construction
On this page

Ask about this page

KEVOS AIReporting a Statistic You Did Not Compute

KEVOS knowledge first · trusted web sources when needed

KEVOS/Project Delivery/Research Projects/Examined Thesis Corpus
Project DeliveryResearch ProjectsAdvancedReporting Results

Reporting a Statistic You Did Not Compute

A reader can do nothing with the word "strong" on its own. Here is what an analysis loses when its apparatus is present and its number is absent, how a single sign confusion produced two wrong results, and the two-minute check that finds both.

Reading time16 minutes
LevelAdvanced
Topic streamReporting Results
Source materialExamined Thesis Corpus
Updated2026-08-16

In brief

  • One examined work introduces correlation, prints the formula as a figure, prints a five-line interpretation key, builds a scoring scheme, produces two correlation plots — and reports no value of R anywhere in the document.
  • Both stated results are described with the wrong sign against the work's own printed key. The error is identical in both, which makes it one error made twice.
  • The scoring scheme itself is a good, reproducible idea. What is missing is everything downstream of it: a coefficient, an n per group, a dispersion measure, and any handling of a described outlier.
  • One of the two relationships is described in the prose as rising and then falling. A linear coefficient is the wrong summary for that shape, whatever its strength.
  • The check is short: every named statistic gets a value; every stated direction gets read back against your own key.

What the document builds, and what it leaves out

The work is a construction-safety study with one online questionnaire and one hundred respondents. Its analysis section names two techniques. The first is thematic analysis, named and never operationalised. The second is correlation, and correlation is where the apparatus is built.

The build is thorough. Correlation is introduced and attributed to a cited statistics source. The formula is printed as a numbered figure. A five-line interpretation key is printed three pages before the results. A bespoke scoring scheme is constructed and its construction is stated exactly. Two scatter plots are produced as numbered figures. Two results are stated in prose, each with a direction and each with a strength.

From the source

The construction, quoted from the work

"A metric was created specifically to quantify the safety of differing work groups on construction sites. To establish the safety index, the average of all participants Likert-scale answers were taken and assigned a value; Strongly Agree = 2, Agree = 1, Neutral = 0, Disagree = -1, and Strongly Disagree = -2."

This is a single scalar per respondent, built by averaging signed Likert values, which lets groups be compared. The construction is stated tightly enough to be reproduced. Treat it as the transferable part of this example.

What never appears is a coefficient. Not one numeric value of R occurs anywhere in the document. The word strong is the entire quantitative content of both results. A reader who wants to know whether "strong" means 0.8 or 0.3 has nothing to consult, because the key that would decide it is printed and the number it would be applied to is not.

The key, and the two results judged against it

The interpretation key is worth reproducing on its own account. It is the clearest short statement of correlation anywhere in this library's material, and it comes from student work rather than from teaching material. The work paraphrases it from a cited statistics source.

THE INTERPRETATION KEY AS PRINTED IN THE WORK (source example, quoted from a cited statistics text)

The key's lineWhat it establishes
"R = a number between -1 and 1"The bound. Any reported value outside it is an error.
"R > 0 = positive association"The sign rule for one direction.
"R < 0 = negative association"The sign rule for the other.
"R = close to 0 indicates a weak linear relationship"The strength rule at one end.
"R = close to -1 or 1 indicates a strong linear relationship"The strength rule at the other.

Quoted from the work, which attributes it to a cited statistics source. It is a source example of how to state a key, not a threshold table: it supplies no cut-off values.

Now the two results. The first concerns age: "The result saw a strong negative association whereby as the construction workers get older their level of safety awareness and compliance increases." The second concerns industry experience: "this plot also revealed a strong negative association where increased time spent in the construction industry equated to increased safety behaviours."

Caution

Both sentences describe a positive association and label it negative

Each sentence has one variable rising as the other rises. By the key printed three pages earlier, that is R greater than zero — a positive association. Both are called negative.

The wording of the error is identical in both places, which is the useful part: this is one error made twice, not two independent slips. A single sign confusion, made once and copied, produced two wrong results. Nothing in the document catches it, because there is no coefficient whose sign could have contradicted the word.

What the absent coefficient actually costs

It is tempting to treat the missing number as a presentational omission — the analysis happened, the number just did not get typed. Treat it instead as a question about what the reader can do with the result. Six things become impossible at once.

  • The strength claim cannot be tested. "Strong" is asserted against a key that defines strength numerically, and the number is absent.
  • The sign claim cannot be checked. A printed coefficient of the wrong sign would have contradicted the sentence beside it, in the way a chart's percentage label contradicts a mistyped percentage elsewhere in this corpus.
  • The base is unknown. No n per group is reported, so a strong association among a handful of respondents and one among ninety cannot be distinguished.
  • No uncertainty is reported. No p value, no interval, no dispersion measure appears anywhere in the document.
  • The scoring scheme cannot be audited. The index averages every Likert answer on a fixed +2 to −2 scale, and at least two of the sixteen Likert statements are negatively worded. Whether they were reverse-scored before averaging is not stated. If they were not, the index is internally inconsistent.
  • Replication is closed off. A reader with the same instrument and a comparable sample has nothing to compare a result to.
Source gap

What this work does not supply, and must not be filled in

The document contains no coefficient, no p value, no n per group, no confidence statement and no reverse-scoring statement. None of these can be reconstructed from anything in the material, and no reader should assume a value.

The same document states no limitations of its own study anywhere — there is no limitations heading, paragraph or sentence in more than fourteen thousand words. The absence of a coefficient is therefore never acknowledged as a limitation either.

A shape a linear coefficient cannot describe

The second result carries a further problem, and the work states it itself. Describing the experience plot, the prose reads: workers in their first six months "tend to be relatively unsafe"; as time in industry increases the safety score improves; "the trend of strong safety behaviours continues for the first 10 years … however, as the tenure continues workers were identified as returning to unsafe practices as complacency sets in."

That is a curve that rises and then falls. A linear correlation coefficient measures how well a straight line fits; a relationship shaped like an arch can have a coefficient near zero however strong and however real the pattern is. The work describes the shape correctly and then reports it as a strong linear association. Both statements are in the same section.

Practice note

Look at the scatter before you choose the statistic

The source does not prescribe this. In practice: plot the two variables first, then choose the summary. If the cloud turns — rises then falls, or falls then rises — a linear coefficient is answering a question you are not asking. Report the shape, or split the range at the turn and describe each segment, or use a measure that does not assume linearity. State which you did.

The same section describes an age band as "a standout outlier" and then explains it by reference to external literature. An explained outlier is still in the correlation. If you keep it, say so and report the result with and without it; if you drop it, say so and say why.

Where the apparatus stops matching the analysis

Three smaller defects sit alongside the missing coefficient, and each is worth recognising because each is easy to make and easy to catch.

THREE APPARATUS DEFECTS IN ONE ANALYSIS SECTION (findings of one examined work)

What the document doesWhy it mattersWhat would have caught it
Introduces the correlation formula as the method "to display linear regression"Correlation and regression are different procedures answering different questions; the cited statistics source distinguishes them.Reading the definition back against the technique you actually ran.
Applies thematic analysis to "the qualitative and quantitative methodologies", and never names a theme, a code or a coding frameA named technique with no procedure and no output cannot be assessed or repeated.For each named technique, point at the paragraph that reports its output.
Directs the reader to "Figure 3" for the safety index by employment position; the caption below reads Figure 5, and Figure 3 is a different chartA cross-reference that is off by two sends a checking reader to the wrong evidence.A final pass reading every figure reference against every caption.

All three are recorded in one work's analysis and discussion sections. They are that document's defects, not a pattern claimed about examined work generally.

Keep the index, fix the reporting

It would be easy to read this example as a case against constructing your own measure. It is the opposite. The index is the strongest thing in the analysis: a defensible way of turning sixteen Likert statements into one comparable score per respondent, with its construction stated so exactly that a reader could rebuild it from the sentence. Very little student work in this corpus states a derived measure that precisely.

Two things would have made it usable as evidence rather than as illustration. First, a statement of how the negatively worded items were handled before averaging. Second, the numbers the index was built to produce — a coefficient per plot, an n per group, and a dispersion measure. Neither is a large piece of work. Both are the difference between a described pattern and a measured one.

Describing a pattern — legitimate, if labelled

  • "Safety scores rise with age across the sample."
  • "The plot rises to about ten years' tenure and falls after it."
  • "The youngest band sits well below the trend."
  • No coefficient is claimed, so none is owed. The reader knows what they are being given.

Reporting a statistic — owes a number

  • "A strong negative association was found."
  • "R indicates a strong linear relationship."
  • Either sentence commits you to a value, a direction that matches it, and a base.
  • If you cannot supply those three, write the description instead.

The pre-flight check

Six passes over your own results section

  1. List every statistic you name

    Search the document for each technique you introduce — correlation, regression, a test, an index, a rate. For each one, point at the paragraph that reports its output. A named technique with no output is either an unreported analysis or a leftover from a plan.

  2. Give every named statistic a value

    A formula, a scoring scheme and an interpretation key are apparatus, not evidence. If you print the key, print the number the key is for, with its n. This is the single check that would have changed the example on this page.

  3. Read your direction words back against your own key

    Take each sentence that states an association, write out which variable rises and which falls, and apply your printed key to it. Do this after the sentence is written, not while writing it — the whole point is that the two were composed at different moments.

  4. Check the shape before you accept the coefficient

    Look at the scatter. If it turns, say so, and either report the turn or choose a summary that survives it. A near-zero coefficient on an arch is a true statement about linearity and a misleading statement about the relationship.

  5. State your scoring decisions

    If you built a composite measure, state the scale, the direction of each item, and whether any item was reverse-scored. One sentence. Without it a reader cannot tell whether a low score means low safety or low agreement with a negatively worded item.

  6. Reconcile figure references to captions

    Read every in-text figure number against the caption it points at, in one pass, at the end. Cross-references drift when figures are inserted or reordered, and drift by a fixed offset — here, by two.

Check before you proceed

Before you submit

Search your document for the names of every statistical technique you mention. For each hit, ask: is there a number attached to this within two paragraphs? If not, either compute and report it, or delete the apparatus and write the description instead. Both are honest; only the middle position — apparatus without a number — leaves the reader unable to judge anything you have said.

What this example does and does not establish

This is one examined work, and it passed. The finding is about the document, not the researcher: a document that sets up a statistic and does not report it, and that states two directions against its own printed key. It is not evidence about how often that happens in examined work generally, and no count here supports such a claim.

It is also worth being clear about why the defect is visible at all. It is visible because the work printed its key. Had the interpretation key been left out, there would be no standard in the document against which the two sentences could be called wrong. The same pattern runs through the whole of this corpus: tracing a figure to its source works because a raw file survived, and when a thesis disagrees with itself works because both accounts were printed. Publishing your apparatus makes your work checkable, and being checkable is not the same as being wrong.

If you are choosing between correlation and another approach, correlation and experimental methodologies sets out what each is for. If you are deciding what belongs in the results chapter at all, writing the results section and descriptive and inferential statistics are the adjacent pages. For what happens when the conclusion then travels further than the statistic ever went, see overreach in conclusions.

What to carry forward

  1. Apparatus is not evidence. A formula, a scoring scheme, an interpretation key and a plot can all be present while the statistic is absent — and a reader can do nothing with the word "strong" on its own.
  2. If you print a key, print the number the key applies to, with its n. If you cannot, write a description of the pattern instead and do not use statistical vocabulary for it.
  3. Read every direction word back against your own key after the sentence is written. One sign confusion, copied, produced two wrong results here.
  4. Plot before you choose the summary. A relationship that rises and then falls will not be captured by a linear coefficient, however strong it looks.
  5. State how a composite measure was scored, including whether negatively worded items were reversed. One sentence closes a hole nobody else can fill.
  6. This is one examined work's document, not a claim about examined work. It is visible only because the work printed the key it then contradicted.

Frequently asked questions

Is it always wrong to describe a relationship without reporting a coefficient?

No. Describing a pattern you can see in a plot is legitimate, provided you describe it as a pattern. The problem arises when statistical vocabulary is used — "strong", "association", "correlation" — because that vocabulary commits you to a measured quantity a reader will expect to see. Choose one register and stay in it.

What should I report alongside a correlation coefficient?

At minimum the value, the number of cases it was computed on, and what the two variables are in the units you measured them. Where your analysis supports it, add a measure of uncertainty. The example on this page reports none of these, which is why the result cannot be assessed even though the plots exist.

How do I know whether a linear coefficient is the right statistic?

Look at the scatter first. If the points trend consistently in one direction, a linear coefficient summarises them usefully. If they rise and then fall, or scatter in bands, the coefficient will be small regardless of how clear the pattern is, and reporting it as weak would misdescribe your data. Describe the shape instead, or analyse the segments separately.

Does building my own index count as a valid measure?

It can, and the one on this page is a reasonable piece of construction. What makes a derived measure usable is that its construction is stated exactly enough for someone else to rebuild it — the scale, the item set, the direction of each item, and the treatment of any negatively worded item. State all four and the measure travels; leave one out and it does not.

What is reverse-scoring, and why does it matter here?

In a Likert battery, some statements are worded so that agreement indicates the opposite of what agreement indicates elsewhere. Before those items are averaged with the others, their scores must be flipped. The index on this page averages all sixteen statements on a fixed scale, at least two of which are negatively worded, and the document does not say whether they were flipped. Without that sentence a reader cannot interpret the index at all.

Does one missing coefficient mean the study is worthless?

No. The same work carries a hundred responses, a fully displayed instrument, an eleven-element consent text and seventy-five verbatim free-text accounts of unsafe work — the largest block of primary qualitative data in this corpus. The defect is confined to how one analysis was reported. Separating a reporting defect from a data defect is exactly the discipline this page is about.

References and source attribution

  1. An examined master's work in project management, supplied as student work: a construction-safety questionnaire study of one hundred respondents whose analysis section introduces correlation, prints an interpretation key and reports no coefficient. Researcher, supervisor, institution, employer, jurisdiction and industry identifiers scrubbed. Used as observed practice, not as a model answer.
  2. The same work's displayed questionnaire (fifteen numbered items including a sixteen-statement Likert matrix) and its printed appendix of seventy-five verbatim free-text responses, against which the instrument observations on this page were checked item by item.
  3. The consolidated extract of four examined works written to one section template, in which every quantitative claim was checked against the printed evidence and the correlation apparatus above was recorded in full.
  4. The supplied teaching source: weekly study notes, slide decks and assessment activities for a master's-level research methods subject in project management, which names inferential statistics and supplies no inferential procedure. Author, institution and year not stated in the supplied files.

Suggested questions for Ask KEVOS

  • What must I report alongside a correlation coefficient in a thesis?
  • How do I tell whether a linear correlation is the wrong statistic for my data?
  • Should I reverse-score negatively worded Likert items before averaging them?
  • How do I write about a pattern in a scatter plot without claiming a statistic?
  • What is the difference between correlation and regression in a results chapter?
  • How do I check my results section for statistics I set up and never reported?

Related KEVOS knowledge

Correlation and Experimental MethodologiesCore · research methodologyDescriptive and Inferential StatisticsCore · quantitative analysisWriting the Results SectionCore · reporting resultsWhen a Thesis Disagrees With ItselfAdvanced · research exemplarsOverreach in ConclusionsCore · reporting resultsResearch Integrity in Examined WorkAdvanced · research exemplars
KEVOS® · Project Delivery · Research Projects Page KVS-PM-RES-0261 · v1.0.0 · content 2026.08 Last reviewed 2026-08-16

Continue learning

Building an Analysis WorkbookGuide · Research ProjectsNEXT LESSON →Overreach in ConclusionsGuide · Research ProjectsData Screening: A Worked DecisionGuide · Research ProjectsWhat a Project Completion Report RequiresGuide · Research Projects
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®