ID/Program | In No. | Filtered No. (%) | Out No. | Pass Rate | Acc. Pass Rate | Filter example and notes
|
---|
Generate the MEDLINE n-gram set
|
---|
Generate MEDLINE n-gram set | 5,235,873,758
- N=1: 30,229,399
- N=2: 273,853,147
- N=3: 935,932,433
- N=4: 1,740,156,534
- N=5: 2,255,702,245
| 5,212,702,625 (99.56%)
- N=1: 29,206,987 (96.62%)
- N=2: 267,853,047 (97.81%)
- N=3: 927,397,909 (99.09%)
- N=4: 1,734,808,402 (99.69%)
- N=5: 2,253,436,280 (99.90%)
| 23,171,133
- N=1: 1,022,412
- N=2: 6,000,100
- N=3: 8,534,524
- N=4: 5,348,132
- N=5: 2,265,965
| 0.4425% | N/A
| From MEDLINE TI & AB to the MDELINE n-gram set
- filter out n-grams with length > 50
- filter out n-grams with word count < 30
- Calculated by Excel (manualy input In and Out No.)
- Used data from n-gram set (not distilled)
|
Basic operation: Sort nGrams by DC|WC|Terms
|
---|
ID-01- NGramFilter: SortNGramByDcWcTerm
- Param: 1, 01
- Run Time: 1 Min.
| 23,171,133 | 0 | 23,171,133 | 100.0000% | 100.0000% |
- Create link: ./05.ApplyFilters/nGram.${YEAR}
|
Apply General Exclusive Filters
|
---|
ID-10
| 23,171,133 | 13 | 23,171,120 | 99.9999% | 99.9999% |
- |
- (|r|
- ||
- Ag|AgCl
- |D|
- |E|
- lambda(||)
|
ID-11
| 23,171,120 | 581 | 23,170,539 | 99.9975% | 99.9974% |
- =
- <
- +/-
- >
- -
- -->
- (+)
- (%)
- "+"
- ((-/-))
- ==>
- [...]
- *}
- *//
|
ID-12
- Filter: Digit
- InTerm: core-term.lc
- Param: 2, 12
- Run time: 2 Min (norm - strip punc and space)
| 23,170,539 | 143,714 | 23,026,825 | 99.3798% | 99.3772% |
- 2
- 1
- 3
- 10
- 4
- 95%
- 2,
- 2000
- 3-5
- +/-0.5
- (+/-0.05)
- $1,500
- "3 + 1"
- 55834
- 192.168.1.1
- [192, 168]
- (+15%),
|
ID-13
- Filter: Number
- InTerm: core-term.lc
- Create link: ./inData/NRVAR
- Param: 2, 13
- Run time: 2 Min
| 23,026,825 | 4,737 | 23,022,088 | 99.9794% | 99.3568% |
- and
- two
- one
- first
- three
- first and second
- one third
- twenty-eight
- NINE
- zeroth and
- Four hundred and forty-seven
- zero-one
- 'half'
- One"
|
ID-14
| 23,022,088 | 177,174 | 22,844,914 | 99.2304% | 98.5921% |
- of the
- in the
- to the
- and the
- on the
- In the
- and/or
- 50% of
- 1, 2, and
- 2003 to
- 2003 to 2007
- for >=50%
- the 8:2
- -196 to -174
- OR-462
- AND-34
- IN-1130
- And-1
|
Apply Exclusive Filters - pattern
|
---|
ID-20
| 22,844,914 | 256,870 | 22,588,044 | 98.8756% | 97.4836% |
- tomography (CT)
- imaging (MRI)
- resonance imaging (MRI)
- oxide (NO)
- reaction (PCR)
- chain reaction (PCR)
- polymerase chain reaction (PCR)
- magnetic resonance imaging (MRI)
- computed tomography (CT)
- enzyme-linked immunosorbent assay (ELISA)
- single nucleotide polymorphisms (SNPs)
- magnetic resonance (MR) imaging
- "Standards, Options and Recommendations" (SOR)
- (CREB)-binding protein (CBP)
- kinase (ASK)
- proline-rich polypeptide (PRP)
- semi-permeable membrane devices (SPMDs)
|
ID-21
| 22,588,044 | 411,712 | 22,176,332 | 98.1773% | 95.7067% |
- a significant
- a single
- a high
- a novel
- a case
- a very
- a group
- a dose-dependent
- A series
- A and B
- a meta-analysis
- a SIF
- A alpha C
- A nonseminomatous
- a delivery rate per
- A beta 2m
- a beta ab
|
ID-22
| 22,176,332 | 146,394 | 22,029,938 | 99.3399% | 95.0749% |
- RESULTS:
- METHODS:
- CONCLUSIONS:
- CONCLUSION:
- BACKGROUND:
- OBJECTIVE:
- OBJECTIVE: To
- MATERIALS AND METHODS:
- SETTING:
- PURPOSE: To
- INTRODUCTION:
- AIM: The
- L: -DOPA
- 95% PI:
- PHPT:
- months [95% CI:
- vs N:
- mode MIC:
- [95 % CI:
|
ID-23
| 22,029,938 | 159,141 | 21,870,797 | 99.2776% | 94.3881% |
- (n =
- (P <
- (P =
- (p <
- P <
- (P < 0.05)
- 95% CI =
- P<0.001),
- CI},
- US$
- VSL#3
- N^N
- group (n=6) received
- CYP3A7*1C
- studies; average
- n.; Trichoteleia
- sp. n.; Trichoteleia
|
ID-24
| 21,870,797 | 373,552 | 21,497,245 | 98.2920% | 92.7760% |
- two groups
- 6 months
- 24 h
- (ABSTRACT TRUNCATED AT 250 WORDS)
- the two groups
- 5 years
- at 37 degrees
- 3 times
- 100 mg
- January 1,
- 10 mg/kg
- 12-year-old
- at -20 degrees C
- September 2006
- 65 years or older with
- 20 cigarettes per day
- 3 - 6 months
- 6 hours plus
- minutes) per day, 5 days
- MMR + V
- 3 mg/EE
- 317615 x
|
ID-25
| 21,497,245 | 196,791 | 21,300,454 | 99.0846% | 91.9267% |
- group (P
- significant (P
- years) with
- significantly (P
- years) and
- interval [95%
- see text] The
- lt; 0.05) lower
- CENTRAL) (The
- nM (SD
- pOGH (ANG
- cB72.3(gamma
- new species (type
|
Apply Exclusive Filters - Lead-End-Terms
|
---|
ID-30
| 21,300,454 | 5,603,439 | 15,697,015 | 73.6933% | 67.7438% |
- of a
- that the
- from the
- is a
- of this
- The results
- was observed
- this study was
- about 50%
- - but not
- "what is
- AND COURSE
- iT reg
- of FoxM1b
- or spinal or conduction
- or spinal or conduction block,
|
ID-31
| 15,697,015 | 3,241,729 | 12,455,286 | 79.3481% | 53.7535% |
- patients with
- associated with
- at the
- suggest that
- between the
- in patients with
- results suggest that
- MATERIALS AND
- cross-reacted with
- (ST 36) and
- Zusanli (ST 36) and
- determine whether this could
- primarily composed of the
- tilt-in-space and
- systems, assays and
- ppm Cu as
- epidural or spinal or
|
ID-32
| 12,455,286 | 3,045 | 12,452,241 | 99.9756% | 53.7403% |
- in a
- to be
- with a
- as a
- may be
- In a
- in A.
- For one
- on NO
- anti-NOR
- plus AT
- I/a
- AS-ON
- anti-OF
|
ID-33
| 12,452,241 | 1,732,833 | 10,719,408 | 86.0842% | 46.2619% |
- to determine
- In addition,
- to evaluate
- to assess
- to investigate
- in the presence
- AT 250
- As a result,
- ON THE TREATMENT
- as a possible treatment for
- in details,
- - for example,
- within working memory
- for various chronic
- in 0.1% trifluoroacetic
- in threatened preterm labor
- with the MIC90S
- On PTD
- plus LHRH-A
- with the MIC90S of
|
ID-34
| 10,719,408 | 1,732,833 | 8,979,895 | 83.7723% | 38.7547% |
- effects of
- number of
- use of
- presence of
- used to
- Comparison of
- low cost of
- HPV) in
- NUMBER OF
- zymography was used to
- loss of two or more
- 1 goes to
- active with the MIC90s of
- syn. nov. of
- microg/mmol of
|
The final results of above is used as the distilled MEDLINE n-gram set
|
---|
Apply Exclusive Filters - Project domain
|
---|
ID-40
| 8,979,895 | 816,410 | 8,163,485 | 90.9085% | 35.2313% |
- of
- the
- in
- to
- a
- The
- We
- "The
- linear,
- "normal"
- {systematic name:
- systematic name
- anterior intermeniscal ligament
- regional low-flow perfusion
- Neo.
- Cannon &
- Polycentropus
- Penneys &
- % (month
|