What to do now
Correct an OCR-extracted equipment identifier by locating the exact source image and text region, preserving the raw extraction, and recording a separately reviewed value with its basis. Distinguish a clear transcription error from an unreadable label, a formatting change or a conflicting physical asset. Recheck any dependent search, asset link or work request before treating the corrected field as ready for use.
Key takeaways
- A reviewed field must point back to the actual source version and region.
- Raw extraction, transcription correction and equipment interpretation are different records.
- A higher OCR score or matching catalog result does not by itself prove the physical asset’s identity.
Define what the field is meant to contain
Start by naming the intended field: manufacturer name, model identifier, serial identifier, internal asset number or another label. These are not interchangeable. A correct reading placed in the wrong field can be as misleading as a misread character. Locate the image and surrounding label text before deciding that the characters need correction.
This workflow concerns record quality. It does not authorize electrical work, opening equipment, entering a unit or selecting a replacement part. Use existing access and qualified-maintenance procedures if a new observation is required. A photo that can be obtained only by unsafe access is not a reason for an unqualified operator to attempt it.
Preserve a usable source reference
Keep the original document or image version, extraction run identity, relevant page or image location, and the raw extracted string. If the extraction tool provides a region or bounding shape, retain that reference with the run. If it does not, record a clear description of where the field appears so another reviewer can find the same evidence.
Do not rely on a bare extraction block ID across different runs. Amazon Textract’s Block reference describes its Id as unique only within a single operation and separately supplies text, geometry and confidence fields. Other providers can use different identifiers, so check their documented scope rather than assuming a block number is a permanent document coordinate.
| Record | Example to preserve | Why it matters |
|---|---|---|
| Source version | Original image attachment and its version | The reviewer needs the same pixels, not a later replacement |
| Source region | Label area, page and bounding region if available | Connects the value to a particular location |
| Extraction run | Run identity plus provider block identity where supplied | Separates two extractions of the same file |
| Raw output | Characters returned before manual correction | Shows the actual error instead of rewriting model output |
| Reviewed value | Approved transcription or unresolved status | Keeps the usable record separate from the raw reading |
| Correction basis | What was inspected, by whom and when | Lets a later operator evaluate the change |
Classify the mismatch before editing
Open enough context to tell whether the extraction selected the intended label. Then decide what sort of discrepancy you have. A clearly visible letter read as a digit can support a transcription correction. A blurred character supports an unresolved reading, not a confident guess. A catalog match may suggest a question to investigate; it does not replace inspection of the correct source.
Keep formatting rules separate. Removing spaces or converting letter case can be a defined normalization without changing the transcription. Substituting an apparent model identifier from another asset is an identity decision, not a formatting fix. Name the kind of change in the correction record so an operator does not interpret all edits as equally supported.
| Observed mismatch | Appropriate record action | Avoid |
|---|---|---|
| Clear source character differs from OCR | Record a reviewed transcription and the exact source region | Overwriting the saved raw extraction |
| Character cannot be read reliably | Keep the field unresolved and request suitable evidence | Choosing whichever candidate looks familiar |
| Text was taken from a neighboring label | Correct field-to-region association and recheck the value | Changing characters while keeping the wrong source reference |
| Spacing differs but characters match | Apply a documented normalization separately if needed | Describing normalization as a newly verified identifier |
| Photo and asset record appear to describe different equipment | Route the asset identity conflict for verification | Silently relinking the source to the convenient asset |
Record the correction as a new accountable assertion
Write the prior raw value, reviewed transcription, source reference, reason, correcting person and correction time. If a second check is required by the team’s policy, record it separately rather than assigning a generic “verified” badge. The scope can be narrow: the reviewer confirmed the characters visible in this source, not the equipment’s current installation or suitability.
Where the source is unreadable, use an explicit unresolved value or a held field through the available workflow. Do not put an unsupported candidate into the active field merely to make an import succeed. If a draft candidate must be retained for investigation, keep it visibly provisional and away from automatic downstream selection.
Recheck the uses of the former value
List what the old field already influenced: search results, asset matching, a draft work request, a vendor question, a maintenance schedule or a parts inquiry. Resolve those dependencies at the level actually affected. Correcting a searchable label may be enough for an unused draft; a request already sent to a contractor needs a separate clarification through the established contact path.
Do not treat the new string as approval for a technical action. A catalog search can return a plausible product while leaving manufacturer, variant, installation and compatibility questions unresolved. The authorized qualified person should verify any equipment or parts decision; this article addresses the information handed to that person.
| Use of the old field | Recheck | Completion evidence |
|---|---|---|
| Unsent draft asset record | Does the corrected field retain its source link? | Reviewed value and lineage visible in the draft |
| Search or duplicate match | Was a prior match based on the old string? | Match decision reevaluated without silently merging assets |
| Vendor inquiry | Did the vendor receive the old identifier? | Corrected inquiry and acknowledgment where needed |
| Scheduled task | Was the schedule selected from the inferred equipment type? | Applicable owner confirms the schedule’s basis |
| Exported inventory | Does the saved version contain the superseded value? | Corrected version or marked follow-up under the export process |
Keep extraction reruns from erasing reviewed corrections
Before reprocessing a document, identify which active fields already have reviewed values. Treat the new extraction as another candidate observation. Compare it with the preserved source and correction record instead of automatically replacing the active value. If the source image itself changed, record the new version and decide which earlier review still applies.
A disagreement between two extraction runs is useful diagnostic information, but it is not a majority vote about the underlying text. Route material disagreements to a reviewer and retain the run identities. Repeated agreement is also not proof that the field came from the correct label or the correct asset.
Compare a rerun at the field and source level
A rerun brings a new observation, not an automatic instruction to replace the reviewed field. Put the old raw output, reviewed transcription and new raw output beside the source version and region that produced each. The first question is whether the runs describe the same evidence. Only then compare the characters and the earlier reviewer’s reason.
An unchanged source with a repeated OCR mistake need not undo a supported manual correction. A changed image may require a new review even when the returned text agrees. Keep the old review attached to what it actually inspected; do not extend it to a new photograph merely because the filename or equipment record stayed the same.
Give each comparison a specific disposition
The following decisions concern transcription evidence. Where a difference suggests another physical asset, stop the field-replacement decision and route the identity question. Neither a more recent run nor a higher confidence value is a universal tie-breaker.
| Source and region | Comparison | Disposition | Evidence to preserve |
|---|---|---|---|
| Same saved image and label | New raw text repeats the old misreading | Retain supported reviewed transcription; record repeated disagreement | Both run IDs, old raw string and review basis |
| Same saved image and label | New raw text agrees with reviewed value | Record agreement without claiming independent physical confirmation | New run and same source reference |
| Same image, different selected label | New text matches neighboring inventory sticker | Correct region selection before evaluating the field | Both regions and their surrounding label text |
| New image version | New text differs from reviewed value | Review the new source and asset identity separately | Old review, new image and unresolved identity question |
| New image version | New text is unchanged | Check whether the earlier review’s scope covers the new source | Version difference and new review decision where needed |
| Different operations reuse a block label | Only block ID appears to match | Do not treat the ID match as source equivalence | Operation identity plus image/page/region references |
Accept the new evidence without losing the earlier decision
Before closing the comparison, confirm that the current field has a documented disposition, each retained candidate has its source/run reference, and any affected downstream request has an owner. Record why an earlier correction was retained, superseded or left pending. A later reviewer should be able to recover that reason without inferring it from the newest timestamp.
Amazon Textract documents that its block ID is scoped to one operation; the comparison therefore keeps operation and source context together. W3C provenance concepts help describe the relationships between observations and revisions. Neither source supplies an automatic conflict-resolution rule, and these examples do not imply that Aptoria uses that OCR provider or implements these controls.
Close the data-quality task without overstating the result
Confirm that the source opens, the region is identifiable, the raw output remains available under the records policy, and the reviewed value is distinguishable from it. Check the dependency register and leave unresolved equipment questions assigned. A closing statement should say which field was corrected and which uses were rechecked.
W3C PROV-O’s derivation and revision concepts provide general vocabulary for source relationships. The field worksheet here is original operational analysis. It does not imply a particular OCR provider is in use, that Aptoria preserves these fields automatically, or that a source-linked value is technically correct for every downstream purpose.
Operational checklist
Mark your progress, then save a working copy. Selections reset when you leave this page. A checked box is not an approval or evidence of completion.
☐
Identify the intended field and correct source label.
☐
Preserve source version, region, raw output and extraction run.
☐
Classify transcription, normalization, wrong-region or asset-identity issues.
☐
Record a supported reviewed value or an explicit unresolved state.
☐
Keep the reviewer’s scope and correction time separate from source capture.
☐
Recheck prior searches, matches, requests and exports that used the old field.
☐
Protect reviewed values from blind extraction reruns.
☐
Close only the corrected field and completed dependencies.
0 of 8 marked
Edge cases
- The original source was replaced in place: recover a versioned source or mark the earlier extraction’s provenance incomplete.
- Several labels contain similar strings: retain the field name and surrounding label, not only a tight character crop.
- A catalog normalizes punctuation differently: keep the source transcription and normalized lookup string separate.
- A new photograph contradicts the old asset record: route the identity conflict rather than declaring the earlier transcription wrong.
Sources and references
Follow each source to check the underlying claim. Access checks and professional review are different steps.
1. Primary source · Amazon Web Services
Amazon Textract API: BlockBlock documents recognized text, geometry, confidence and identifiers. Its Id is unique only within one operation, so a block identifier alone is not a cross-run source reference.
Source checked 2026-09-06
Automated source-access check: 2026-09-06.
2. Primary source · World Wide Web Consortium
PROV-O: The PROV OntologyPROV-O models entities, activities, responsible agents and derivation or revision links. It supplies provenance vocabulary, not proof that a maintenance observation is true.
Published 2013-04-30 · Source checked 2026-09-06
Automated source-access check: 2026-09-06.
Continue the workflow
property management field mapping guideChoose the source of truth for property operationsCorrect Maintenance Photos Attached to the Wrong UnitRevision history
2026-09-06
Initial source-linked OCR field-correction workflow with explicit AI-assisted technical review.
2026-09-06
Expanded rerun reconciliation with a six-case matrix and fictional multi-run/source packet preserving the scope of prior reviewed values.