Showing posts with label OCR. Show all posts
Showing posts with label OCR. Show all posts

Tuesday, April 10, 2012

Phylogeny digitisation

I was hunting around for further research on phylogeny image digitisation to see whether any advances had been made since I last published on the topic and to keep my previous post up-to-date. The main reason behind all of this is to see whether there would be a faster way to digitise a bunch of images that are accumulating on my hard drive. I thought it would be cool to do something with the ripped phylogenies for the iEvoBio Challenge but my current set of trees only has a total of 2,000 leaves and I need 10,000.
Anyway, I came across PHYLODIGM in my searches which looks promising. Thomas Laubach has also done some further work on TreeSnatcher Plus including using the benchmarking dataset from TreeRipper and a number of tree files found via Google searches. Additionally, he has released the source code under the GNU General Public License.
This all looks promising!!!

Sunday, May 22, 2011

Recognition of tree images



I have just published a program on the automated recognition of phylogenies from tree images. First of all, I would like to apologise for the use of the word 'towards' in the title, I know that it is increasingly being used and irritating to some. I just wanted to be honest in that this program does not succeed on all tree images but is a step in the right direction.

I thought I would take the opportunity to post a few links to software that deals with the same problem. We have all been rather unorginal with names!

To my knowledge, this was the first program that dealt with the problem of converting a phylogenetic image back to the more useful bracket format such as NEXUS or newick. It requires the user to click on tips and nodes in a specific order and type in the label at the tips. Unfortunately, this program only works on MacOS 9.

TreeSnatcher was a conceptual advance on TreeThief, relies heavily on Java libraries and is cross-platform. It requires a limited amount of input from the user, such as selecting the foreground and background and lets the user improve the quality of the extraction thanks to this interactivity.

TreeSnatcher Plus is an improvement on TreeSnatcher as it lets the user convert almost anything to a newick file, for example it works on radial tree images.

TreeRogue is essentially the same concept as TreeThief and I have just come across this so unfortunately do not make reference to it in my paper (sorry). It uses an R script that converts coordinates to a tree file. These coordinates can be detected from an image by using GraphClick, which costs $8.

TreeRipper has been written in C++ and there is a version running on the website, the code is available under GNU GPL v3. It uses heavily the C++ API to to ImageMagick image-processing library (Magick++) and it uses Tessecract-ocr to convert the leaf labels to text. This is a fully automated approach that unfortunately only works on a proportion of the tree images. You could for example use TreeRipper for batch processing a large number of trees and then use a semi-automated approach for the trees that weren't converted.

There is still a lot of room for improvement and I am hoping that someone out there will make further progress on this interesting challenge.

Of course, none of these programs would be necessary if we all shared our trees and this would only be possible if we had a useful phylogenetic standard <= this statement should please the TDWG Interest Group on Phylogenetic Standards ;)

Reference:
Hughes, J. (2011). TreeRipper web application: towards a fully automated optical tree recognition software. BMC Bioinformatics 12: 178 doi: 10.1186/1471-2105-12-178

Tuesday, July 03, 2007

Installing tesseract command line OCR on MacOS X

Installing libpng from source:
http://kenno.wordpress.com/2006/04/20/compiling-libpng-for-mac-os-x/

fink install libjpeg, aspell, aspell-en

I will want to create my own aspell dictionary using taxonomic names:
http://www.mail-archive.com/code4lib@listserv.nd.edu/msg01545.html

Download and installing tesseract following install instructions:
http://code.google.com/p/tesseract-ocr/downloads/list

fink xpdf for pdfimages to extract images from a pdf:
>pdfimages -j LandPlants_paper.pdf LandPlantImg

To convert in imagemagick to tif for tesseract :
convert LandPlantImg.jpg -compress None test.tif

Using tesseract:
tesseract test.tif out.txt

I have now got a script to extract the names and check them against a dictionary of taxonomic names from spira.
I am thinking that using information from the article itself might provide even better results. When tesseract 2.0 comes out, there will also be a way of training the program to improve the character recognition. OCRupus also looks like an interesting program for layout detection but it doesn't work on MacOSx yet
The line extraction is proving to be much more difficult than first thought mainly because the lack of consistn format and the labelling at the nodes that get in the way of edge detection. I have tried a number of methods for cleaning up the image and bit by bit I will get there, I hope.






Thursday, June 07, 2007

Branch thinning and tracking

I think I might have found a way to convert phylogenetic trees from image to nexus format through a process similar to that used in GIS map vectorization. This approach should also enable me to deal with trichotomies. The pattern matching approach that I had used previously turned out to be unsuccessful because of the many inconsistencies in tree drawing making it difficult to find a pattern that would always match a tip or a node. Line tracking offers hope!
All this would be so unnecessary if only researcher submitted their phylogenies to TREEBASE. Still, there are a number of phylogenies published prior to Treebase.

Wednesday, March 14, 2007

OCR taxa names from phylogenies

I have been trying out different open source OCR software, to recognise the taxa used in phylogenies. Tesseract (newly released by Google) did not fair as well as GOCR. If you train GOCR using a database of character images, it does even better.
________________________________
Tesseract
________________________________
M mii
is A
' 53 V S
i
3 7 @1
M `W gj i :[
T * 5?$fA
`Gy i i 8
V ~ 'E S` v~v :
;A bg fh , [
A `i $ > ggt wd
E V 1 Jvh xg E 4 i
A 3 ^ Awbt ) 4 V M
Tvlrjg A
W [ M Y 2 At A] [ vp
> 4 ii ; h i1
% V b a i % i i S
4-E A ) * k i VZ-WVYA S 4`
4 W AQAA A [A~Jw Q QT
VV 42 V V i mvi Q. A A4@ $
@ N [ 2 h Jii@ ' S
,1 > ) ~~%# A -` X * @8
A T V S E -3 A~I: W i
A ` J V V g - M ? A AE
3 `@ J g t. ,4
1 jji WE jzi Q E >
$ A ag H L 1 @ Q w$
vi V BE * V 3 V ` * miji ,)
I V `V as ` 7 ,! I i
5 ) i AVVM *4.3 ~# 4 p u ` 1 ~ V *
i J .41 " gm ^ ` V` U N9 7* 6) *`
2 =
` ~ " ,4 ` ` ` L.;i A>?, ps > ,` i V. g>v..Ve,g $?i[ X ` I j:)4;w . ( 'X A " K 7
4 1 qvb~: JJ S t *%%`~
T - V (A %%^) I As
3 I jwi @ @ A W J [
?g M 6* i Pg`( h b I S W
1 i ; v ^ T [f L AA A
v A H 2 3 Y44@ M V V [ A M { Vi
5 ? El t 52 Wi 1* AL i
* 1^ 4@=b> i-sm ^ 44J
Q *i b g N %A J
4 ; V A C A i ]i T J L ^ A$v J
1 ,V;gt t izm dmgg 4 A 7
T ~ "as A ` . , @ J AVE V7;z`* I?
i i O gm ] is Aig
`- ` V An 7 F x ``^ w ` `An pb _ A ~ e h 7 U ^,
W L - N 4 A #14$A i
M A @ * i
V Z V gi A E T v?)
` pig [A$[ X A 11 44 i L V
~$ Ivd? A ^~hr4 V mg
i $ ihb 4 4 A A p LLAE
& V L Ay M #
N A if '4jv[ 1 x
` V.>gg( A QA VIV \ v4Ji`
J7* **nif if @;ibA Mb 2 W A i & $@^@ Q
% & L% J 4i @ @w@ 4%
1 A x ][A Q ii ] ijN i&i A [
V A " F
i L t WAA! ia G iLAi*
!^=p~ .4 x vvgv V A
I - G 24 @1 V `'A. p m
VAF A i
1 @A[f A 2 X X 4A i @
W gv@ I gpa 'A @ ' E
,>gn I At M i-44 ` BE 4
m m M I V OF J; { V A V qil
)~'; E` > FF a ;; M `* ga ~
LA v 6 @ YAVZ4 A V 3 ) M VY J 8 (
A! i h A 1`3A& X *
HE % At _ ` 7 @ 2 " ' 5 4 "
S 9 4,p; 5
k A V j , i v [ MY >yAA SS 4
V ` ^%jV 4wJv` J 4 4
` Q? W ppgawi V 1 . 1
\ 1 _` 7 2* I-
74 [ ^*^- Av pw
2 ~T$ 'W`` V 1 i i i
Q) H pi In as i
3 4 1 'AY @4 ~ 4, Y
1 V Er A [ M $95 V 8 i `
V ( @4 V * ) jg.: A
V Q 4 s 1 ( vv ;
44 g E 3 iii A] >;1A 9;
C; 4%%3 2 P 4`4A%yV;` 4s
` v4'Y H4:tA? K ; i4 VA 1 VAV A`=$ g `
=?h V) V
4 s iqvv i @M,ib J .l@ A S V?
wbi` i4 $; ~ Q? V t
Aii 6 4 2; EYE V )AGv
A W g 1 ) S HE Jig
1 % AM fA4LN ii O Egg?
V; V @3 ;#
` V I iv V7>:v2 V VA N
* 'id; E ' X? V %.`i:h; 'M( @6i . 7 _ ` t V} .5 '
`V V4 gvgv E1`

M Q`;`A
' 'Apq
~ $$i
___________________________________
GOCR
___________________________________
c(PICTURE)100 C. _asicus (14)
C. nasicus (_8)
C. jo_e_sjs (30)
C. co_fusoy (32)
C. coMfusoy (12)
C. su_cafu_us (_3)
C. su_cafu_us (68)
C. humey__is(_4)
C. vicoyiensis (_5)
C. p_y__lis (26)
C. vicfoyie_sÌs (54)
C. _ongi_e_s (16)
C. longi_e_s (19)
C. ve_osus (82)
C. veMos_s (166)
C. ve_osus (160)
C. ve_osus (149)
C. s__icivoyus
C.pe_lifus (131)
C. pellifus (151)
C. pellifus (192)
C. eleph_s (_18)
C. eleph_s (_13)
C. eleph_s (_16)
C. c___e (204)
C. cR_Re (205)
C. ca__e (208)
C. pyobosci_eus (58)
C. humey__is (56)
C. p_o_osci_eus (22)
C. scufel1_pjs (5)
C. nucu_ (117)
C. nucum (_43)
C. g____ium (_OO)
C. g__n_ium (gg)
Pakjsfan _p. (_1)
C. c__elli_e (49)
C. c__el_i_e (20)
__iica_ sp. (8)
C. p__yhoce_as (1)
C. p_y_hoceyas (16)
_______________________
GOCR using database option
_______________________
c(PICTURE)100 C. nasicus (14)
C. nasicus (28)
C. iowensis (30)
C. confusor (32)
C. coMfusor (12)
C. sulcatulus (23)
C. sulcatulus (68)
C. humeyalis(24)
C. vicoyiensis (25)
C. paydalis (26)
C. vicforiensÌs (54)
C. longidens (16)
C. longidens (19)
C. venosus (82)
C. veMosus (166)
C. venosus (160)
C. venosus (149)
C. salicivoyus
C.pellifus (131)
C. pellitus (151)
C. pellitus (192)
C. elephas (218)
C. elephas (213)
C. elephas (216)
C. cawae (204)
C. cawae (205)
C. cawae (208)
C. pyoboscideus (58)
C. humeyalis (56)
C. proboscideus (22)
C. scufel1aris (5)
C. nucum (117)
C. nucum (243)
C. glandium (200)
C. glandium (99)
Pakistan sp. (21)
C. camelliae (49)
C. camelliae (20)
Afiican sp. (8)
C. pyrrhoceras (1)
C. pyyrhoceras (16)

Disqus for Evo-Karma