Efficiently identifying top k similar entities

Hanasoge Sudheendra, Supreetha

Efficiently identifying top k similar entities

Services

Deutsch English

About the Repository Search and Browse Publish

Home
→
Fakultäten
→
Fakultät für Elektrotechnik und Informatik
→
View Item

Download statistics - Document (COUNTER):

Hanasoge Sudheendra, Supreetha: Efficiently identifying top k similar entities. Hannover : Gottfried Wilhelm Leibniz Universität, Master Thesis, 2020, 81 S. DOI: https://doi.org/10.15488/10466

Selected time period:

Sum total of downloads: 241

distribution of downloads over the selected time period
downloads by country

back to single item view (close usage statistics)

FileEfficiently_ident ...

Size3.17 MB

FormatAdobe PDF

View

Abstract:
With the rapid growth in genomic studies, more and more successful researches are being produced that integrate tools and technologies from interdisciplinary sciences. Computational biology or bioinformatics is one such field that successfully applies computational tools to capture and transcribe biological data. Specifically in genomic studies, detection and analysis of co-occurring mutations is an leading area of study. Concurrently, in the recent years, computer science and information technology have seen an increased interest in the area association analysis and co-occurrence computation. The traditional method of finding top similar entities involves examining every possible pair of entities, which leads to a prohibitive quadratic time complexity. Most of the existing approaches also require a similarity measure and threshold beforehand to retrieve the top similar entities. These parameters are not always easy to tune. Heuristically, an adaptive method can have wider applications for identifying the top most similar pair of mutations (or entities in general). In this thesis, we have presented an algorithm to efficiently identify top k similar pair of mutations using co-occurrence as the similarity measure. Our approach used an upperbound condition to iteratively prune the search space and tackled the quadratic complexity. The empirical evaluations show that the proposed approach shows the computational efficiency in terms of execution time and accuracy of our approach particularly in large size datasets. In addition, we also evaluate the impact of various parameters like input size, k on the execution time in top k approaches. This study concludes that systematic pruning of the search space using an adaptive threshold condition optimizes the process of identifying top similar pair of entities.
License of this version:	Es gilt deutsches Urheberrecht. Das Dokument darf zum eigenen Gebrauch kostenfrei genutzt, aber nicht im Internet bereitgestellt oder an Außenstehende weitergegeben werden.
Document Type:	MasterThesis
Publishing status:	publishedVersion
Issue Date:	2020-12-28
Appears in Collections:	Fakultät für Elektrotechnik und Informatik

distribution of downloads over the selected time period:

downloads by country:

pos.	country		downloads
pos.	country		total	perc.
1		Germany	92	38.17%
2		United States	39	16.18%
3		China	17	7.05%
4		No geo information available	14	5.81%
5		Russian Federation	13	5.39%
6		United Kingdom	6	2.49%
7		India	5	2.07%
8		Israel	5	2.07%
9		Czech Republic	5	2.07%
10		Austria	5	2.07%
		other countries	40	16.60%

Further download figures and rankings:

Hinweis

Zur Erhebung der Downloadstatistiken kommen entsprechend dem „COUNTER Code of Practice for e-Resources“ international anerkannte Regeln und Normen zur Anwendung. COUNTER ist eine internationale Non-Profit-Organisation, in der Bibliotheksverbände, Datenbankanbieter und Verlage gemeinsam an Standards zur Erhebung, Speicherung und Verarbeitung von Nutzungsdaten elektronischer Ressourcen arbeiten, welche so Objektivität und Vergleichbarkeit gewährleisten sollen. Es werden hierbei ausschließlich Zugriffe auf die entsprechenden Volltexte ausgewertet, keine Aufrufe der Website an sich.

Search the repository

Browse

All content
- Communities & Collections
- By Issue Date
- Authors
- Titles
- Subjects
- Subjects (GND)
- DDC
- License
- Type
This Collection
- By Issue Date
- Authors
- Titles
- Subjects
- Subjects (GND)
- DDC
- License
- Type

Efficiently identifying top k similar entities

Download statistics - Document (COUNTER):

Selected time period:

Sum total of downloads: 241

distribution of downloads over the selected time period:

downloads by country:

Further download figures and rankings:

Search the repository

Browse

All content

This Collection