TomasFAV commited on
Commit
b93719d
·
verified ·
1 Parent(s): 959e0b4

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +92 -20
README.md CHANGED
@@ -1,41 +1,110 @@
1
  ---
2
  library_name: transformers
 
 
3
  tags:
4
  - generated_from_trainer
 
 
 
 
 
5
  metrics:
6
  - precision
7
  - recall
8
  - f1
9
  - accuracy
10
  model-index:
11
- - name: BERTInvoiceCzechRV2
12
  results: []
13
  ---
14
 
15
- <!-- This model card has been generated automatically according to the information the Trainer had access to. You
16
- should probably proofread and complete it, then remove this comment. -->
17
 
18
- # BERTInvoiceCzechRV2
19
 
20
- This model was trained from scratch on an unknown dataset.
21
  It achieves the following results on the evaluation set:
22
- - Loss: 0.1326
23
- - Precision: 0.8120
24
- - Recall: 0.7868
25
- - F1: 0.7992
26
- - Accuracy: 0.9700
 
 
27
 
28
  ## Model description
29
 
30
- More information needed
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
31
 
32
- ## Intended uses & limitations
33
 
34
- More information needed
 
 
 
35
 
36
- ## Training and evaluation data
 
 
37
 
38
- More information needed
 
 
 
 
 
39
 
40
  ## Training procedure
41
 
@@ -52,6 +121,8 @@ The following hyperparameters were used during training:
52
  - num_epochs: 10
53
  - mixed_precision_training: Native AMP
54
 
 
 
55
  ### Training results
56
 
57
  | Training Loss | Epoch | Step | Validation Loss | Precision | Recall | F1 | Accuracy |
@@ -67,10 +138,11 @@ The following hyperparameters were used during training:
67
  | 0.0733 | 9.0 | 783 | 0.1385 | 0.7893 | 0.7930 | 0.7912 | 0.9686 |
68
  | 0.0733 | 10.0 | 870 | 0.1393 | 0.8044 | 0.7938 | 0.7991 | 0.9696 |
69
 
 
70
 
71
- ### Framework versions
72
 
73
- - Transformers 5.0.0
74
- - Pytorch 2.10.0+cu128
75
- - Datasets 4.0.0
76
- - Tokenizers 0.22.2
 
1
  ---
2
  library_name: transformers
3
+ license: apache-2.0
4
+ base_model: google-bert/bert-base-multilingual-cased
5
  tags:
6
  - generated_from_trainer
7
+ - invoice-processing
8
+ - information-extraction
9
+ - czech-language
10
+ - synthetic-data
11
+ - hybrid-data
12
  metrics:
13
  - precision
14
  - recall
15
  - f1
16
  - accuracy
17
  model-index:
18
+ - name: BERTInvoiceCzechR-V2
19
  results: []
20
  ---
21
 
22
+ # BERTInvoiceCzechR (V2 – Synthetic + Random Layout + Real Layout Injection)
 
23
 
24
+ This model is a fine-tuned version of [google-bert/bert-base-multilingual-cased](https://huggingface.co/google-bert/bert-base-multilingual-cased) for structured information extraction from Czech invoices.
25
 
 
26
  It achieves the following results on the evaluation set:
27
+ - Loss: 0.1326
28
+ - Precision: 0.8120
29
+ - Recall: 0.7868
30
+ - F1: 0.7992
31
+ - Accuracy: 0.9700
32
+
33
+ ---
34
 
35
  ## Model description
36
 
37
+ BERTInvoiceCzechR (V2) represents an advanced stage in the training pipeline, combining synthetic data with realistic document layouts.
38
+
39
+ The model performs token-level classification to extract structured invoice fields:
40
+ - supplier
41
+ - customer
42
+ - invoice number
43
+ - bank details
44
+ - totals
45
+ - dates
46
+
47
+ This version introduces a key improvement: **real invoice layouts with synthetic content**, bridging the gap between artificial and real-world data.
48
+
49
+ ---
50
+
51
+ ## Training data
52
+
53
+ The dataset is composed of three main components:
54
+
55
+ 1. **Synthetic template-based invoices**
56
+ 2. **Synthetic invoices with randomized layouts**
57
+ 3. **Hybrid invoices with real layouts and synthetic content**
58
+
59
+ ### Real layout injection
60
+
61
+ In the hybrid dataset:
62
+ - real invoice documents are used as layout templates
63
+ - original textual content is removed
64
+ - fields (e.g., supplier, customer, bank details) are replaced with synthetic data
65
+ - new content is rendered into the original spatial structure
66
+
67
+ This approach preserves:
68
+ - realistic spacing
69
+ - typography patterns
70
+ - structural complexity
71
+
72
+ while maintaining:
73
+ - full control over annotations
74
+ - label consistency
75
+
76
+ ---
77
+
78
+ ## Role in the pipeline
79
+
80
+ This model corresponds to:
81
+
82
+ **V2 – Synthetic + layout augmentation + real layout injection**
83
+
84
+ It is designed to:
85
+ - reduce the domain gap between synthetic and real invoices
86
+ - evaluate the impact of realistic spatial distributions
87
+ - serve as a bridge between purely synthetic training (V0–V1) and real data fine-tuning (V3)
88
+
89
+ ---
90
 
91
+ ## Intended uses
92
 
93
+ - Advanced research in document AI
94
+ - Evaluation of hybrid synthetic-real training strategies
95
+ - Invoice information extraction in semi-realistic conditions
96
+ - Benchmarking generalization improvements
97
 
98
+ ---
99
+
100
+ ## Limitations
101
 
102
+ - Still does not use fully real textual content
103
+ - Synthetic text may not capture all linguistic variability
104
+ - OCR noise and scanning artifacts are not fully represented
105
+ - Performance may still drop on unseen real-world edge cases
106
+
107
+ ---
108
 
109
  ## Training procedure
110
 
 
121
  - num_epochs: 10
122
  - mixed_precision_training: Native AMP
123
 
124
+ ---
125
+
126
  ### Training results
127
 
128
  | Training Loss | Epoch | Step | Validation Loss | Precision | Recall | F1 | Accuracy |
 
138
  | 0.0733 | 9.0 | 783 | 0.1385 | 0.7893 | 0.7930 | 0.7912 | 0.9686 |
139
  | 0.0733 | 10.0 | 870 | 0.1393 | 0.8044 | 0.7938 | 0.7991 | 0.9696 |
140
 
141
+ ---
142
 
143
+ ## Framework versions
144
 
145
+ - Transformers 5.0.0
146
+ - PyTorch 2.10.0+cu128
147
+ - Datasets 4.0.0
148
+ - Tokenizers 0.22.2