4
4.1RatingtheQualityofDatabasesNecessaryProceduresforGoodnessEstimation
Theamountofdatainpracticaldatabasesisoftenlarge.Tocomputetheexactsoundnessandcompletenessofaparticularviewwewouldneedto(1)authenticateeveryvaluepairinthestoredview,and(2)determinehowmanypairsaremissingfromthisview.Thismethodisclearlyinfeasibleinanyrealsystem.Thus,wemustresorttosamplingtechniques[16,4].Samplingtechniquesallowustoestimatethemeanandvarianceofaparticularparameterofapopulationbyusingasamplewhichisusuallyonlyafractionofthesizeoftheentirepopulation.Thetheoryofstatisticsalsogivesusmethodsforestablishingasamplesizetoachievepredeterminedaccuracyoftheestimates.Itisthenpossibletosupplementourestimateswithcon denceintervals.Formoredetaileddiscussiononsamplingfromdatabasesthereaderisreferredtotheliteratureonthetopic(see,forexample,[12]foragoodsurvey).Notethattwodi erentpopulationsmustbesampled.Toestimatesoundnesswesamplethegiven(stored)view,whereastoestimatecompleteness,wesampletheidealview.Toestablishbothsoundnessandcompletenessitisnecessarytohaveaccesstotheidealdatabase.Forsoundness,weneedtodeterminewhetheraspeci cvaluepairofthestoreddatabaseisintheidealdatabase.Forcompleteness,itisnecessarytodeterminewhetheraspeci cpairfromtheidealdatabaseisinthestoreddatabase.Theseprocedures(verifyapairfromastoreddatabaseagainsttheidealdatabaseandretrieveanarbitrarypairfromtheidealdatabase)mustbeimplementedinanad-hocmanner[1].Foreachconcretedatabase,humanexpertisewillberequired.Theexpertwillaccessavarietyofavailablesourcestoperformthesetwoprocedures.Notethatthise ortisperformedonlyonceandonlyforasample,whichthenhelpsestimatetheoverallgoodness.
Acriticalstageofoursolutionistobuildasetofhomogeneousviewsonastoreddatabase,calledagoodnessbasis.Thegoodnessoftheviewsofthisbasiswillbemeasuredandthere-afterusedinestablishingthegoodnessofanswerstoarbitraryqueriesagainstthisdatabase.Sincewecannotguaranteeasinglesetofviewsthatwillbehomogeneouswithrespecttobothqualitymeasures,weconstructtwoseparatesets:asoundnessbasisandacompletenessbasis.Inconstructingeachbasis,weconsidereachdatabaserelationindividually.Eachre-lationmaybepartitionedbothhorizontally(byaselection)andvertically(byaprojection),andthebasiscomprisestheunionofallsuchpartitions.Selectionsarelimitedtoranges;i.e.,theselectioncriteriaisaconjunctionofconditions,whereeachindividualconditionspeci esanattributeandarangeofpermittedvaluesforthisattribute.
Weassigntoanincorrectvaluepairthevalue0andtoacorrectpairthevalue1.Thus,wecanrepresentanerrordistributionpatterninaviewextensionasatwo-dimensionalmatrixof0sand1s,inwhichrowscorrespondtothetuplesandcolumnscorrespondtotheattributesoftheview.Avalueinaparticularcellofthismatrixiseither0or1dependingonthecorrectnessofthecorrespondingpairofattributevalues.Wecallthisnewdatastructure
搜索“diyifanwen.net”或“第一范文网”即可找到本站免费阅读全部范文。收藏本站方便下次阅读,第一范文网,提供最新人文社科Estimating the quality of data in relational databases(8)全文阅读和word下载服务。
相关推荐: