【Vue】vue3 在图片上渲染 OCR 识别后的文本框、可复制文本组件
一、背景与问题
在图像处理场景中,我们常常需要将 OCR(光学字符识别)识别后的文本内容以可视化的方式叠加在原始图片上,同时提供可复制的交互功能。这种需求常见于以下场景:
- 文档数字化:将纸质文档中的文字提取后标注在图片上
- 图像标注工具:开发图像标注平台时的辅助功能
- 图像内容分析:在图像中高亮显示识别出的关键信息
传统实现方式存在以下问题:
- 文本位置不精确导致视觉污染
- 无法实现双击复制等交互功能
- 未考虑图片缩放时的坐标映射问题
- 未实现文本框的样式控制(如边框、背景色等)
二、基本原理
本方案主要包含三个核心环节:
- OCR识别:使用开源OCR库(如Tesseract.js)对图片进行文字识别
- 文本框绘制:将识别结果转化为矩形框并叠加在图片上
- 交互功能:实现文本框的复制、编辑、样式控制等交互
关键点在于:
- 坐标映射:确保OCR识别的坐标与图片的像素比例一致
- 动态渲染:使用Vue的响应式特性实现实时更新
- 事件处理:为文本框添加点击、复制等交互事件
三、环境准备
1. 技术栈
- 前端:Vue3 + TypeScript
- OCR库:Tesseract.js(开源)
- 图像处理:Canvas API
- 额外依赖:
@types/tesseract.js(类型定义)
2. 安装依赖
npm install tesseract.js
npm install @types/tesseract.js --save-dev四、核心实现
1. OCR识别模块
// src/services/ocrService.ts
import { Tesseract, TesseractOptions } from 'tesseract.js';
export interface OCRResult {
text: string;
boundingBoxes: { x: number; y: number; width: number; height: number }[];
}
export async function performOCR(file: File): Promise<OCRResult> {
const img = await loadImage(file);
const options: TesseractOptions = {
logger: (progress) => {
console.log(`Progress: ${progress.status}`);
}
};
const { data: { text, boundingBoxes } } = await Tesseract.recognize(
img,
'eng',
options
);
return { text, boundingBoxes };
}关键点解释:
- 使用
loadImage函数将文件转为Canvas对象 boundingBoxes包含每个文字块的坐标信息- Tesseract.js的识别结果包含文本内容和坐标信息
2. 文本框渲染组件
<!-- src/components/OCRTextOverlay.vue -->
<template>
<div class="image-container">
<img :src="imageSrc" ref="imageRef" @click="handleImageClick" />
<div
v-for="(box, index) in textBoxes"
:key="index"
:style="getTextBoxStyle(box)"
class="text-box"
>
<div class="text-content">{{ box.text }}</div>
</div>
</div>
</template>
<script lang="ts">
import { ref, onMounted, onBeforeUnmount } from 'vue';
import { performOCR } from '@/services/ocrService';
export default {
props: {
imageSrc: {
type: String,
required: true
}
},
setup(props) {
const imageRef = ref<HTMLImageElement | null>(null);
const textBoxes = ref<OCRResult['boundingBoxes']>([]);
const isDragging = ref(false);
const dragStart = ref({ x: 0, y: 0 });
const selectedBox = ref<number | null>(null);
const getTextBoxStyle = (box: any) => {
return {
position: 'absolute',
left: `${box.x}px`,
top: `${box.y}px`,
width: `${box.width}px`,
height: `${box.height}px`,
border: '2px solid #00f',
borderRadius: '4px',
padding: '4px',
cursor: 'move'
};
};
const handleImageClick = async (event: MouseEvent) => {
if (!imageRef.value) return;
const rect = imageRef.value.getBoundingClientRect();
const x = event.clientX - rect.left;
const y = event.clientY - rect.top;
// 模拟点击识别文本框
const clickedBox = textBoxes.value.find(box =>
x >= box.x && x <= box.x + box.width &&
y >= box.y && y <= box.y + box.height
);
if (clickedBox) {
selectedBox.value = textBoxes.value.indexOf(clickedBox);
}
};
const handleDragStart = (event: MouseEvent, index: number) => {
isDragging.value = true;
dragStart.value = { x: event.clientX, y: event.clientY };
selectedBox.value = index;
};
const handleDragEnd = () => {
isDragging.value = false;
};
const handleDrag = (event: MouseEvent) => {
if (!isDragging.value || selectedBox.value === null) return;
const dx = event.clientX - dragStart.value.x;
const dy = event.clientY - dragStart.value.y;
const newBox = { ...textBoxes.value[selectedBox.value] };
newBox.x += dx;
newBox.y += dy;
textBoxes.value.splice(selectedBox.value, 1, newBox);
dragStart.value = { x: event.clientX, y: event.clientY };
};
onMounted(async () => {
const result = await performOCR(new File([''], 'temp.jpg', { type: 'image/jpeg' }));
textBoxes.value = result.boundingBoxes;
});
onBeforeUnmount(() => {
document.removeEventListener('mousemove', handleDrag);
document.removeEventListener('mouseup', handleDragEnd);
});
return {
imageRef,
textBoxes,
getTextBoxStyle,
handleImageClick,
handleDragStart,
handleDragEnd,
handleDrag
};
}
};
</script>
<style scoped>
.image-container {
position: relative;
width: 100%;
max-width: 800px;
border: 1px solid #ccc;
}
.text-box {
position: absolute;
border: 2px solid #00f;
border-radius: 4px;
padding: 4px;
cursor: move;
}
.text-content {
white-space: pre-wrap;
word-wrap: break-word;
}
</style>关键点解释:
- 使用绝对定位将文本框叠加在图片上
- 通过
getTextBoxStyle动态计算样式 - 实现文本框的拖拽移动功能
- 使用Vue的响应式特性实时更新文本框位置
3. 可复制文本组件
<!-- src/components/TextCopyButton.vue -->
<template>
<button
@click="copyText"
class="copy-button"
>
<span>复制</span>
</button>
</template>
<script lang="ts">
import { ref } from 'vue';
export default {
props: {
text: {
type: String,
required: true
}
},
setup(props) {
const copyText = () => {
navigator.clipboard.writeText(props.text)
.then(() => {
alert('复制成功');
})
.catch((err) => {
console.error('复制失败:', err);
alert('复制失败');
});
};
return {
copyText
};
}
};
</script>
<style scoped>
.copy-button {
position: absolute;
top: 10px;
right: 10px;
padding: 8px 12px;
background-color: #00f;
color: white;
border: none;
border-radius: 4px;
cursor: pointer;
}
</style>五、完整案例
1. 综合组件实现
<!-- src/components/OCRImageViewer.vue -->
<template>
<div class="image-viewer">
<label for="imageUpload" class="upload-label">上传图片</label>
<input
id="imageUpload"
type="file"
accept="image/*"
@change="handleImageUpload"
/>
<div class="image-container" v-if="imageSrc">
<img :src="imageSrc" ref="imageRef" @click="handleImageClick" />
<div
v-for="(box, index) in textBoxes"
:key="index"
:style="getTextBoxStyle(box)"
class="text-box"
>
<div class="text-content">{{ box.text }}</div>
<TextCopyButton :text="box.text" />
</div>
</div>
</div>
</template>
<script lang="ts">
import { ref, onMounted, onBeforeUnmount } from 'vue';
import { performOCR } from '@/services/ocrService';
import TextCopyButton from './TextCopyButton.vue';
export default {
components: { TextCopyButton },
props: {
imageSrc: {
type: String,
required: true
}
},
setup(props) {
const imageRef = ref<HTMLImageElement | null>(null);
const textBoxes = ref<OCRResult['boundingBoxes']>([]);
const isDragging = ref(false);
const dragStart = ref({ x: 0, y: 0 });
const selectedBox = ref<number | null>(null);
const getTextBoxStyle = (box: any) => {
return {
position: 'absolute',
left: `${box.x}px`,
top: `${box.y}px`,
width: `${box.width}px`,
height: `${box.height}px`,
border: '2px solid #00f',
borderRadius: '4px',
padding: '4px',
cursor: 'move'
};
};
const handleImageClick = async (event: MouseEvent) => {
if (!imageRef.value) return;
const rect = imageRef.value.getBoundingClientRect();
const x = event.clientX - rect.left;
const y = event.clientY - rect.top;
// 模拟点击识别文本框
const clickedBox = textBoxes.value.find(box =>
x >= box.x && x <= box.x + box.width &&
y >= box.y && y <= box.y + box.height
);
if (clickedBox) {
selectedBox.value = textBoxes.value.indexOf(clickedBox);
}
};
const handleDragStart = (event: MouseEvent, index: number) => {
isDragging.value = true;
dragStart.value = { x: event.clientX, y: event.clientY };
selectedBox.value = index;
};
const handleDragEnd = () => {
isDragging.value = false;
};
const handleDrag = (event: MouseEvent) => {
if (!isDragging.value || selectedBox.value === null) return;
const dx = event.clientX - dragStart.value.x;
const dy = event.clientY - dragStart.value.y;
const newBox = { ...textBoxes.value[selectedBox.value] };
newBox.x += dx;
newBox.y += dy;
textBoxes.value.splice(selectedBox.value, 1, newBox);
dragStart.value = { x: event.clientX, y: event.clientY };
};
const handleImageUpload = (event: Event) => {
const file = (event.target as HTMLInputElement).files?.[0];
if (!file) return;
const reader = new FileReader();
reader.onload = (e) => {
const img = new Image();
img.onload = () => {
props.imageSrc = img.src;
performOCR(file).then(result => {
textBoxes.value = result.boundingBoxes;
});
};
img.src = e.target?.result as string;
};
reader.readAsDataURL(file);
};
onMounted(() => {
document.addEventListener('mousemove', handleDrag);
document.addEventListener('mouseup', handleDragEnd);
});
onBeforeUnmount(() => {
document.removeEventListener('mousemove', handleDrag);
document.removeEventListener('mouseup', handleDragEnd);
});
return {
imageRef,
textBoxes,
getTextBoxStyle,
handleImageClick,
handleDragStart,
handleDragEnd,
handleDrag,
handleImageUpload
};
}
};
</script>
<style scoped>
.image-viewer {
position: relative;
width: 100%;
max-width: 800px;
margin: 20px auto;
}
.upload-label {
display: block;
margin-bottom: 10px;
cursor: pointer;
}
.image-container {
position: relative;
width: 100%;
}
.text-box {
position: absolute;
border: 2px solid #00f;
border-radius: 4px;
padding: 4px;
cursor: move;
}
.text-content {
white-space: pre-wrap;
word-wrap: break-word;
}
</style>2. 使用示例
<!-- src/App.vue -->
<template>
<div id="app">
<OCRImageViewer imageSrc="https://via.placeholder.com/800x600" />
</div>
</template>
<script lang="ts">
import { defineComponent } from 'vue';
import OCRImageViewer from './components/OCRImageViewer.vue';
export default defineComponent({
components: { OCRImageViewer },
setup() {
return {};
}
});
</script>六、源码解析
1. OCR识别流程
// Tesseract.js内部处理流程
// 1. 将图片转为Canvas
// 2. 使用Tesseract.js进行文字识别
// 3. 返回包含文字内容和坐标信息的结构关键点:
- Tesseract.js使用的是基于CNN的深度学习模型
- 默认语言为英文('eng')
- 可以通过
addLanguage添加其他语言支持
2. 坐标映射处理
// 在Canvas中处理图片时,需要考虑缩放比例
const scale = image.width / originalWidth;3. 事件处理机制
// 使用Vue的事件系统绑定交互
@mousemove
@mouseup七、进阶使用
1. 动态样式控制
<template>
<div :style="getTextBoxStyle(box, isSelected)">
...
</div>
</template>
<script>
const getTextBoxStyle = (box, isSelected) => {
return {
...baseStyle,
border: isSelected ? '2px solid red' : '2px solid #00f',
backgroundColor: isSelected ? '#f0f0f0' : 'transparent'
};
};
</script>2. 多语言支持
import { addLanguage } from 'tesseract.js';
addLanguage('chi_sim', 'chi_sim.traineddata');3. 精度优化
const options: TesseractOptions = {
// 增加精度参数
config: {
tessedit_pagesegmode: '3'
}
};八、性能与工程实践
1. 性能优化策略
| 优化策略 | 说明 |
|---|---|
| 使用Web Worker | 避免阻塞主线程 |
| 图片压缩 | 降低图片尺寸加快识别 |
| 缓存识别结果 | 避免重复识别相同图片 |
| 异步处理 | 避免页面卡顿 |
2. 安全注意事项
- 避免直接处理用户上传的文件
- 使用沙箱环境运行OCR识别
- 对识别结果进行内容过滤
- 对敏感信息进行脱敏处理
3. 异常处理
try {
const result = await performOCR(file);
} catch (error) {
console.error('OCR识别失败:', error);
alert('识别过程中出现错误,请重试');
}九、常见问题与踩坑
1. 坐标映射错误
问题:文本框位置不准确
原因:未考虑图片缩放比例
解决方案:
const scale = image.width / originalWidth;
boxes.forEach(box => {
box.x *= scale;
box.y *= scale;
});2. 文本框重叠
问题:多个文本框重叠显示
解决方案:使用z-index控制层级
.text-box {
position: absolute;
z-index: 1000;
}3. 复制功能失效
问题:复制按钮点击无响应
原因:未正确绑定事件
解决方案:确保事件监听器正确绑定
<TextCopyButton
:text="box.text"
@copy="handleCopy"
/>4. 性能瓶颈
问题:大图片识别速度慢
解决方案:使用Web Worker进行异步处理
// worker.js
self.onmessage = (event) => {
const { imageData } = event.data;
Tesseract.recognize(imageData, 'eng', { ... })
.then(result => self.postMessage(result));
};十、最佳实践
| 场景 | 推荐方案 |
|---|---|
| 快速标注 | 使用Tesseract.js+Canvas实现 |
| 高精度需求 | 使用百度OCR+后端处理 |
| 大规模处理 | 使用Web Worker+队列机制 |
| 安全敏感场景 | 使用沙箱环境+内容过滤 |
| 多语言支持 | 使用Tesseract.js语言包 |
十一、总结
本方案通过结合Vue3的响应式特性、Tesseract.js的OCR能力,实现了在图片上动态渲染识别文本框并支持复制功能的完整解决方案。该方案适用于需要图像标注、内容识别等场景,但需要注意以下几点:
- 适用场景:适合需要快速实现图像标注、文本识别的中小型项目
- 性能限制:大尺寸图片识别可能影响性能,建议进行优化处理
- 安全风险:需注意用户隐私数据的处理和存储
- 精度限制:OCR识别结果受图片质量、文字排版等因素影响
实际开发中,建议根据具体需求选择合适的OCR库(如需要高精度可选择商业OCR服务),并配合前端组件实现更丰富的交互功能。通过合理的设计和优化,可以构建出高效的图像识别标注系统。