| Example #1 | He had to repeatedly traverse the inner palace’s chambers, yet each attempt ended in failure. |
| Testing Points | “failure”. |
| OmniAlign | 他得不断地一再地穿过内宫里的屋子;可是他一直没有成功。 |
| OmniAlign w/o S.1 | 他得不断地一再地穿过内宫里的屋子;可是他一直没有成功。 |
| Example #2 | 已知有血液疾病及尿酸性肾结石的患者不推荐使用本品,二岁以下儿童不得服用。 |
| Testing Points | 有; 二岁 |
| OmniAlign | This product is not recommended for patients with known blood disorders or uric acid kidney stones, and it should not be taken by children under the age of two. |
| BinaryAlign | This product is not recommended for patients with known blood disorders or uric acid kidney stones, and it should not be taken by children under the age of two. |
| Example #3 | I recently returned to d2 after several year, now I’m windering: Where do you guys sell/buy your stuff? Do you just make a game “O xxx N yyy” and hope for the best? Or are there a website that’s more efficient? |
| Testing Points | returned; windering |
| OmniAlign | 时隔多年,我又 重新开始玩《暗黑破坏神2》(DiabloII),现在我 想知道:大家都是在哪里进行物品交易的?是通过自己创建名为“OxxxNyyy”的游戏房间来交易,然后听天由命吗?还是有更高效的交易网站? |
| BinaryAlign | 时隔多年,我又重新开始玩《暗黑破坏神2》(DiabloII),现在我想知道:大家都是在哪里进行物品交易的?是通过自己创建名为“OxxxNyyy”的游戏房间来交易,然后听天由命吗?还是有更高效的交易网站? |
| Example #4 | 在技术演进的脉络中,预训练模型的出现无疑是里程碑式的突破。以 BERT、GPT、T5 等为代表的模型通过在海量无标注文本上进行自监督学习,习得语言的深层语义表示和语法结构,极大地提升了下游任务的性能上限。这些模型基于 Transformer 架构的多头注意力机制,能够捕捉文本中的长距离依赖关系和上下文关联,为文本分类、情感分析、机器翻译、问答系统等经典任务提供了强大的通用能力。近年来,多语言预训练模型(如 XLM-R、mT5、LaBSE)进一步打破了语言壁垒,通过对数百种语言的联合训练,实现了跨语言文本的理解与生成,为全球化信息传播和跨文化交流提供了技术支撑。在产业应用层面,NLP 技术已渗透到金融、医疗、教育、传媒等多个领域,催生了一系列创新产品和服务。在金融行业,智能客服系统能够基于用户的自然语言查询快速提供账户咨询、业务办理等服务,平均响应时间缩短至秒级,客户满意度提升 30% 以上;风险控制模型通过分析企业年报、新闻舆情等文本数据,精准识别信用风险和市场波动信号,帮助金融机构降低不良贷款率。在医疗领域,医学文本分析系统可自动提取电子病历中的关键信息(如病症、用药史、检查结果),辅助医生进行诊断决策,减少误诊率;多语言医疗翻译工具则为跨境医疗合作提供了语言保障,使不同国家的医护人员能够高效协作。在教育领域,智能写作辅助系统能够实时检测文本中的语法错误、逻辑漏洞,并提供优化建议,帮助学生提升写作能力;个性化学习平台通过分析学生的学习行为和文本交互数据,精准推送适配的学习资源,实现 “因材施教” 的 教育理念。 |
| Testing Points | (Bert、GPT、T5); (多头注意力机制); (文本分类、情感分析、机器翻译、问答系统); (XLM-R、mT5、LaBSE); (创新产品和服务); (智能客服系统); (风险控制模型); (多语言医疗翻译工具); (实时检测); (学习行为和文本交互数据); (教育理念) |
| OmniAlign | In the context of technological evolution, the emergence of pre-trained models is undoubtedly a milestone breakthrough. Represented by BERT, GPT, T5, and other models, they acquire deep semantic representations and syntactic structures of language through self-supervised learning on massive unlabeled text, greatly raising the performance ceiling of downstream tasks. Based on the multi-head attention mechanism of the Transformer architecture, these models can capture long-distance dependencies and contextual correlations in text, providing powerful general capabilities for classic tasks such as text classification, sentiment analysis, machine translation, and question-answering systems. In recent years, multilingual pre-trained models (e.g., XLM-R, mT5, LaBSE) have further broken down language barriers. Through joint training on hundreds of languages, they have realized the understanding and generation of cross-lingual text, providing technical support for global information dissemination and cross-cultural communication. At the industrial application level, NLP technology has penetrated into multiple fields such as finance, healthcare, education, and media, spawning a series of innovative products and services. In the financial industry, intelligent customer service systems can quickly provide account consultations, business processing, and other services based on users’ natural language queries, reducing the average response time to seconds and increasing customer satisfaction by more than 30%; risk control models accurately identify credit risks and market fluctuation signals by analyzing text data such as corporate annual reports and news public opinion, helping financial institutions reduce the non-performing loan ratio. In the healthcare field, medical text analysis systems can automatically extract key information from electronic medical records (such as symptoms, medication history, and examination results) to assist doctors in diagnostic decisions and reduce the misdiagnosis rate; multilingual medical translation tools provide language guarantees for cross-border medical cooperation, enabling medical staff from different countries to collaborate efficiently. In the education field, intelligent writing assistance systems can real-time detect grammatical errors and logical flaws in text, and provide optimization suggestions to help students improve their writing skills; personalized learning platforms accurately push adaptive learning resources by analyzing students’ learning behaviors and text interaction data, realizing the educational concept of "teaching students in accordance with their aptitude". |
| BinaryAlign | In the context of technological evolution, the emergence of pre-trained models is undoubtedly a milestone breakthrough. Represented by BERT, GPT, T5, and other models, they acquire deep semantic representations and syntactic structures of language through self-supervised learning on massive unlabeled text, greatly raising the performance ceiling of downstream tasks. Based on the multi-head attention mechanism of the Transformer architecture, these models can capture long-distance dependencies and contextual correlations in text, providing powerful general capabilities for classic tasks such as text classification, sentiment analysis, machine translation, and question-answering systems. In recent years, multilingual pre-trained models (e.g., XLM-R, mT5, LaBSE) have further broken down language barriers. Through joint training on hundreds of languages, they have realized the understanding and generation of cross-lingual text, providing technical support for global information dissemination and cross-cultural communication. At the industrial application level, NLP technology has penetrated into multiple fields such as finance, healthcare, education, and media, spawning a series of innovative products and services. In the financial industry, intelligent customer service systems can quickly provide account consultations, business processing, and other services based on users’ natural language queries, reducing the average response time to seconds and increasing customer satisfaction by more than 30%; risk control models accurately identify credit risks and market fluctuation signals by analyzing text data such as corporate annual reports and news public opinion, helping financial institutions reduce the non-performing loan ratio. In the healthcare field, medical text analysis systems can automatically extract key information from electronic medical records (such as symptoms, medication history, and examination results) to assist doctors in diagnostic decisions and reduce the misdiagnosis rate; multilingual medical translation tools provide language guarantees for cross-border medical cooperation, enabling medical staff from different countries to collaborate efficiently. In the education field, intelligent writing assistance systems can real-time detect grammatical errors and logical flaws in text, and provide optimization suggestions to help students improve their writing skills; personalized learning platforms accurately push adaptive learning resources by analyzing students’ learning behaviors and text interaction data, realizing the educational concept of "teaching students in accordance with their aptitude". |
| Scenario 2: Sentence Alignment |
| Example #1 | Source (1) Their Australian born captain, the world’s top-ranked match racer Peter Gilmour, lived in Japan for three years to satisfy the Cup’s crew-nationality rules. (2) Syndicate head Tatsumitsu Yamasaki, chairman of spice giant S&B Foods, is hungry for a win. Target (1) 该船队的船长为世界顶尖帆船赛选手皮得·吉尔摩。 (2) 为了达到杯赛在船员国籍方面的各项要求,这位出生在澳大利亚的选手在日本居住了3年。 (3) 山崎达光是一家财团的总裁兼调味品行业的巨人S&B食品株式会社董事会主席。 (4) 他渴望日本船队能够获得胜利。 |
| OmniAlign | [(1)]:[(1), (2)]; [(2)]:[(3), (4)] |
| BertAlign | [(1)]:[(1), (2)]; [(2)]:[(3)]; []:[(4)] |
| Example #2 | Source (1) Then came the hard part: getting people to want to see the thing. (2) Here’s how they did it, from whisper to buzz to big box-office noise, in only 21 steps. Target (1) 随后进入最艰苦的宣传炒作阶段。 (2) 下面就是他们的宣传炒作步骤:刚开始知者甚少,随后观众渐增,最后票房炙手可热。 (3) 这一成绩的取得仅需21步:。 |
| OmniAlign | [(1)]:[(1), (2)]; [(2)]:[(2), (3)] |
| BertAlign | [(1), (2)]:[(1), (2), (3)] |
| Example #3 | Source (1) The situation in Japan has to change. " (2) At Nissan, it already has. Target (1) 这种形势必须改变。 (2) 在日产,这种变革已经开始。 |
| OmniAlign | [(1)]:[(1)]; [(2)]:[(2)] |
| BertAlign | [(1), (2)]:[(1), (2)] |
| Example #4 | Source (1) 就在这种情况下,我想起十五队的队医陈清扬是北医大毕业的大夫,对针头和勾针大概还能分清,所以我去找她看病。 (2) 看完病回来,不到半个小时,她就追到我屋里来,要我证明她不是破鞋。 Target (1) Under the circumstances, I recalled that the doctor at the fifteenth team, Chen Qingyang, had graduated from Beijing Medical School. (2) Maybe she would be able to tell the difference between a hypodermic and a crotchet needle. (3) So I went to see her. (4) Not half an hour after my visit, she chased after me to my room, wanting me to prove that she wasn’t damaged goods. |
| OmniAlign | [(1)]:[(1), (2), (3)]; [(2)]:[(4)] |
| BertAlign | [(1)]:[(1), (2)]; [(2)]:[(3), (4)] |