The MedDRA version you froze is part of your results
An adverse event table looks like a count of things that happened. It is a count of things that happened, as classified by a dictionary that is revised twice a year, applied by people exercising judgement, at a version somebody chose.
None of that makes the table wrong. It does mean the version and the coding conventions are part of the result, and they are frequently the part nobody wrote down.
Five levels, and the counting happens in the middle
The hierarchy runs from lowest level term at the bottom, through preferred term, high level term and high level group term, up to system organ class. Verbatim text is coded to a lowest level term. Almost every table anyone reads is counted at preferred term, grouped by system organ class.
The two ends of that chain behave differently. Lowest level terms are stable and numerous, because they exist to absorb the many ways people describe the same thing. System organ classes are stable and few. The middle, where the counting happens, is where revision activity concentrates, and it is the level whose changes move numbers.
It is also worth remembering that a preferred term belongs to more than one system organ class, with one designated primary. Tables are built on the primary assignment. Secondary links exist and are useful for signal work, and a reviewer asking about them is not confused.
What a version change actually does
Two releases a year, each carrying additions, changes and, less often, promotions and demotions between levels. In a single study most of it is irrelevant. The parts that are not fall into three groups.
New preferred terms appear, and events that previously fell under a broader term move onto a more specific one. The total is unchanged; the rows split. A table that showed one moderately common event now shows two uncommon ones, and a reader comparing across studies sees a reduction that did not happen.
Terms are made non-current, and their events have to be recoded somewhere. Where they land is decided by the recoding, not by the dictionary.
Primary system organ class assignments change. This one moves whole blocks of a table without changing any preferred term count, and it is the least visible of the three.
Freezing and upversioning are both defensible. Silence is not
Coding a study at one version and holding it there gives internal consistency, which is what matters for the study's own analysis. The cost is that events which arrived late are coded against a dictionary that no longer reflects current terminology, and that pooling this study with a later one requires work.
Upversioning gives comparability across a programme, which is what matters for an integrated summary. The cost is that previously reviewed tables change, sometimes after they have been discussed, and every change needs an explanation.
Both positions are reasonable. What causes trouble is a programme where each study made the decision independently, nobody recorded it, and the integrated summary discovers three versions in use across five studies with no mapping between them.
Decide the version policy for the programme, not for the study, and decide it before the first study locks. It is a five minute conversation at the start and a month of reconciliation at the end.
The coder is making a judgement, and it should be a consistent one
Auto-encoders match verbatim text against a synonym list and resolve most events without help. The remainder go to a person, and that is where consistency is won or lost.
Two events reported as chest discomfort and chest pain may be the same clinical event described by two investigators, or they may not be. Coding them to the same preferred term merges a signal; coding them separately splits one. Neither choice is available to an algorithm without a convention, and the convention is the sponsor's to set.
The practical control is a coding guideline maintained per programme, not per study, covering the terms your therapeutic area actually generates. It should say what to do with combined events, with events reported as worsening of an existing condition, and with symptoms that are part of the indication under study. That last case is the one that produces the most inconsistency and the most reviewer questions.
Record enough to reconstruct the table
Four things, kept with the data rather than in a memo.
The dictionary version, per study and per data cut, in the datasets and in define.xml. A study whose safety data were coded across two versions because of a mid-study upversion needs both recorded, with the point of change.
The verbatim text, preserved unchanged. It is the only thing that survives every version and the only thing a recoding exercise can work from.
The coding path, meaning the lowest level term as well as the preferred term. A table built only from preferred terms cannot be re-derived under a different version without going back to verbatim, and going back to verbatim means recoding rather than mapping.
Whether each code was assigned automatically or manually. When a reviewer questions a grouping, the first useful question is whether a person made that call, and a dataset that cannot answer it turns a short query into an investigation.
None of these are expensive to capture during the study. All of them are expensive to reconstruct afterwards, which is the general shape of every problem in this part of the work.