• Stars
    star
    283
  • Rank 146,066 (Top 3 %)
  • Language
  • License
    Creative Commons ...
  • Created over 4 years ago
  • Updated over 1 year ago

Reviews

There are no reviews yet. Be the first to send feedback to the community and the maintainers!

Repository Details

KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding

KorNLU Datasets

This is the dataset repository for our paper KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding.

We introduce KorNLI and KorSTS, which are NLI and STS datasets in Korean.

KorNLI

Dataset Overview

KorNLI Total Train Dev. Test
Source - SNLI, MNLI XNLI XNLI
Translated by - Machine Human Human
# Examples 950,354 942,854 2,490 5,010
Avg. # words (premise) 13.6 13.6 13.0 13.1
Avg. # words (hypothesis) 7.1 7.2 6.8 6.8

Examples

Example English Translation Label
P: ์ €๋Š”, ๊ทธ๋ƒฅ ์•Œ์•„๋‚ด๋ ค๊ณ  ๊ฑฐ๊ธฐ ์žˆ์—ˆ์–ด์š”.
H: ์ดํ•ดํ•˜๋ ค๊ณ  ๋…ธ๋ ฅํ•˜๊ณ  ์žˆ์—ˆ์–ด์š”.
I was just there just trying to figure it out.
I was trying to understand.
Entailment
P: ์ €๋Š”, ๊ทธ๋ƒฅ ์•Œ์•„๋‚ด๋ ค๊ณ  ๊ฑฐ๊ธฐ ์žˆ์—ˆ์–ด์š”.
H: ๋‚˜๋Š” ์ฒ˜์Œ๋ถ€ํ„ฐ ๊ทธ๊ฒƒ์„ ์ž˜ ์ดํ•ดํ–ˆ๋‹ค.
I was just there just trying to figure it out.
I understood it well from the beginning.
Contradiction
P: ์ €๋Š”, ๊ทธ๋ƒฅ ์•Œ์•„๋‚ด๋ ค๊ณ  ๊ฑฐ๊ธฐ ์žˆ์—ˆ์–ด์š”.
H: ๋‚˜๋Š” ๋ˆ์ด ์–ด๋””๋กœ ๊ฐ”๋Š”์ง€ ์ดํ•ดํ•˜๋ ค๊ณ  ํ–ˆ์–ด์š”.
I was just there just trying to figure it out.
I was trying to understand where the money went.
Neutral

KorSTS

Dataset Overview

KorSTS Total Train Dev. Test
Source - STS-B STS-B STS-B
Translated by - Machine Human Human
# Examples 8,628 5,749 1,500 1,379
Avg. # words 7.7 7.5 8.7 7.6

Examples

Example English Translation Label
ํ•œ ๋‚จ์ž๊ฐ€ ์Œ์‹์„ ๋จน๊ณ  ์žˆ๋‹ค.
ํ•œ ๋‚จ์ž๊ฐ€ ๋ญ”๊ฐ€๋ฅผ ๋จน๊ณ  ์žˆ๋‹ค.
A man is eating food.
A man is eating something.
4.2
ํ•œ ๋น„ํ–‰๊ธฐ๊ฐ€ ์ฐฉ๋ฅ™ํ•˜๊ณ  ์žˆ๋‹ค.
์• ๋‹ˆ๋ฉ”์ด์…˜ํ™”๋œ ๋น„ํ–‰๊ธฐ ํ•˜๋‚˜๊ฐ€ ์ฐฉ๋ฅ™ํ•˜๊ณ  ์žˆ๋‹ค.
A plane is landing.
A animated airplane is landing.
2.8
ํ•œ ์—ฌ์„ฑ์ด ๊ณ ๊ธฐ๋ฅผ ์š”๋ฆฌํ•˜๊ณ  ์žˆ๋‹ค.
ํ•œ ๋‚จ์ž๊ฐ€ ๋งํ•˜๊ณ  ์žˆ๋‹ค.
A woman is cooking meat.
A man is speaking.
0.0

License

Creative Commons Attribution-ShareAlike license (CC BY-SA 4.0)

References

If you use KorNLI or KorSTS for research, please cite our paper:

@article{ham2020kornli,
  title={KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding},
  author={Ham, Jiyeon and Choe, Yo Joong and Park, Kyubyong and Choi, Ilji and Soh, Hyungjoon},
  journal={arXiv preprint arXiv:2004.03289},
  year={2020}
}

More Repositories

1

fast-autoaugment

Official Implementation of 'Fast AutoAugment' in PyTorch.
Python
1,587
star
2

nerf-factory

An awesome PyTorch NeRF library
Python
1,265
star
3

pororo

PORORO: Platform Of neuRal mOdels for natuRal language prOcessing
Python
1,252
star
4

coyo-dataset

COYO-700M: Large-scale Image-Text Pair Dataset
Python
1,062
star
5

kogpt

KakaoBrain KoGPT (Korean Generative Pre-trained Transformer)
Python
1,000
star
6

torchgpipe

A GPipe implementation in PyTorch
Python
776
star
7

karlo

Python
679
star
8

rq-vae-transformer

The official implementation of Autoregressive Image Generation using Residual Quantization (CVPR '22)
Jupyter Notebook
669
star
9

mindall-e

PyTorch implementation of a 1.3B text-to-image generation model trained on 14 million image-text pairs
Python
630
star
10

word2word

Easy-to-use word-to-word translations for 3,564 language pairs.
Python
350
star
11

torchlars

A LARS implementation in PyTorch
Python
326
star
12

g2pm

A Neural Grapheme-to-Phoneme Conversion Package for Mandarin Chinese Based on a New Open Benchmark Dataset
Python
326
star
13

trident

A performance library for machine learning applications.
Python
176
star
14

autoclint

A specially designed light version of Fast AutoAugment
Python
170
star
15

sparse-detr

PyTorch Implementation of Sparse DETR
Python
150
star
16

hotr

Official repository for HOTR: End-to-End Human-Object Interaction Detection with Transformers (CVPR'21, Oral Presentation)
Python
132
star
17

kortok

The code and models for "An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks" (AACL-IJCNLP 2020)
Python
114
star
18

bassl

Python
113
star
19

scrl

PyTorch Implementation of Spatially Consistent Representation Learning(SCRL)
Python
108
star
20

flame

Official implementation of the paper "FLAME: Free-form Language-based Motion Synthesis & Editing"
Python
103
star
21

brain-agent

Brain Agent for Large-Scale and Multi-Task Agent Learning
Python
92
star
22

helo-word

Team Kakao&Brain's Grammatical Error Correction System for the ACL 2019 BEA Shared Task
Python
88
star
23

jejueo

Jejueo Datasets for Machine Translation and Speech Synthesis
Python
74
star
24

solvent

Python
66
star
25

noc

Jupyter Notebook
44
star
26

cxr-clip

Python
43
star
27

expgan

Python
41
star
28

autowu

Official repository for Automated Learning Rate Scheduler for Large-Batch Training (8th ICML Workshop on AutoML)
Python
39
star
29

nvs-adapter

Python
33
star
30

ginr-ipc

The official implementation of Generalizable Implicit Neural Representations with Instance Pattern Composers(CVPRโ€™23 highlight).
Python
30
star
31

coyo-vit

ViT trained on COYO-Labeled-300M dataset
Python
28
star
32

irm-empirical-study

An Empirical Study of Invariant Risk Minimization
Python
28
star
33

coyo-align

ALIGN trained on COYO-dataset
Python
25
star
34

magvlt

The official implementation of MAGVLT: Masked Generative Vision-and-Language Transformer (CVPR'23)
Python
23
star
35

hqtransformer

Locally Hierarchical Auto-Regressive Modeling for Image Generation (HQ-Transformer)
Jupyter Notebook
21
star
36

CheXGPT

Python
18
star
37

learning-loss-for-tta

"Learning Loss for Test-Time Augmentation (NeurIPS 2020)"
Python
9
star
38

stg

Official implementation of Selective Token Generation (COLING'22)
Jupyter Notebook
8
star
39

leco

Official implementation of LECO (NeurIPS'22)
Python
6
star
40

bc-hyperopt-example

brain cloud hyperopt example (mnist)
Python
3
star