188宝金博页面版

  • 图案背景
  • 纯色背景
视图
标记
批注
批注本地保存成功,开通会员云端永久保存 去开通
wangfeifei..

上传于:2015-03-28

粉丝量:7

该文档贡献者很忙,什么也没留下。


  • 相关
  • 目录
  • 笔记
  • 书签

188宝金博页面版:更多相关文档

  • Spam Filtering- Online Naive Bayes Based on TONE

    星级: 14 页

  • spam filtering: online naive bayes based on tone

    星级: 12 页

  • Second Order Online Collaborative Filtering

    星级: 16 页

  • Online filtering of CO

    星级: 9 页

  • 【精品】Study of the online event filtering algorithm for BESⅢ

    星级: 9 页

  • 【精品】Relaxed Online SVMs in the TREC Spam Filtering Track

    星级: 10 页

  • 体外循环中含氧血持续肺动脉灌注的肺保护作用

    星级: 4 页

  • Behavior patterns of online users and the effect on information filtering

    星级: 8 页

  • advances in online learning-based spam filtering

    星级: 212 页

  • towards online spam filtering in social networks

    星级: 16 页

暂无目录

点击鼠标右键菜单,创建目录

暂无笔记

选择文本,点击鼠标右键菜单,添加笔记

暂无书签

在左侧文档中,点击鼠标右键,添加书签

188宝金博页面版: [精品]Distributed Online Filtering

下载积分: 280

内容提示: Distributed Online FilteringCostin RaiciuDavid S. RosenblumUniversity College London{c.raiciu|d.rosenblum|m.handley}@cs.ucl.ac.ukProject URL: http://www.cs.ucl.ac.uk/staff/craiciu/onlinefiltering/Mark HandleyWeb search is a reactive process: search engines in-dex the information publishedon many millions ofservers,and users send queries to find sites that contain partic-ular information. While this approach is adequate forstatic content, it fails to capture the temporal nature ofonline information. Sports ...

文档格式:PDF | 页数:2 | 浏览次数:78 | 上传日期:2015-03-28 12:10:51 | 文档星级:
Distributed Online FilteringCostin RaiciuDavid S. RosenblumUniversity College London{c.raiciu|d.rosenblum|m.handley}@cs.ucl.ac.ukProject URL: http://www.cs.ucl.ac.uk/staff/craiciu/onlinefiltering/Mark HandleyWeb search is a reactive process: search engines in-dex the information publishedon many millions ofservers,and users send queries to find sites that contain partic-ular information. While this approach is adequate forstatic content, it fails to capture the temporal nature ofonline information. Sports scores, networking researchpreprints, stormy blog discussions, and indeed news areall examples of information whose value is strictly re-lated to its timely delivery to users. Additionally, it isa bit awkward that the very numerous web servers pub-lish information and basically hope that the (tens of)thousands ofmachines owned by search companies willmake their content available to the world. Perhaps theweb servers, having spare capacity most ofthe time (ex-cept for the rare case offlash crowds), can eliminate themiddle-man and deliver information to interested usersdirectly?This work focuses on online filtering of documentsby the web servers themselves, organized in an overlay.Users express their long term interests as queries andregister them with the web servers. Content publishedby the web servers is matched against users’ interestsand a decision is made to decide whether a given docu-ment should be delivered to the user. This approach isproactive, as decisions are taken online based on con-tent. It does not come to replace search, but to comple-ment it by enabling timely information delivery to theusers who have expressed interest in that information.Requirements and Metrics. Ourmain metric is through-put, measured as the number of documents processedper second by the whole network. We require the sys-tem should be scalable according to this parameter: byadding more nodes (i.e. web servers) and maintain-ing the same number of user queries, the maximumthroughput should improve.Constraints. Measurements presented elsewhere [1]show that the product between the number of queriesstored and the number of documents processed is al-most constant for a node: this is the node’s computingpower. Further, each node has a limited bandwidth bud-get it can use. Both constraints can be measured; alter-natively, a budget can be specified to avoid degradingthe performance ofthe web server.Document{K1,K2,K3}Query{K1,K4}K1K2K3K4K5(a) Keyword AlgorithmDocumentLoadBalancerLoadBalancerQuery(b) Matrix AlgorithmFigure 1. Basic AlgorithmsSolution. There are two classes of solutions: content-sensitive and content agnostic.Content-sensitive solutions use keywords to route doc-uments and store queries in the overlay. Solutions rangefrom assigning single keywords to nodes to assigningall possible combinations ofkeywords to nodes (in CANstyle). In this space, we create the Keyword algorithmthat assigns every keyword to a single node in DHT-style (Fig. 1.a). Queries are broken up into keywordsand stored on each of the nodes in charge of those key-words.The most important keywords in documents(25-50 keywords) are extracted and used to route thedocument. Keywords in documents andqueries are con-sidered in alphabetical order, such that it is easy to de-termine when a document has been matched against allsignificant queries and also whether a document shouldbe sent to a particular user upon a match. We con-sidered several context-sensitive algorithms, but foundthem to be worse than this simple approach [1].Content-agnostic solutions simply aimto rendez-vouseach document with all the queries. The most generalalgorithm replicates subscriptions R times and routesdocuments to N/R nodes. Nodes are arranged in a ma-trix with R columns and N/R rows, replicating querieshorizontally and routing documents vertically: the Ma-trix algorithm (Fig. 1.b).Comparison. We now analytically compare Matrix andKeyword. For this, we consider a simple theoreticalmodel that assumes nodes are homogeneous and com-pute the maximum throughput without loss of the twoalgorithms. For the Matrix algorithm, we derive thevalue of R that minimizes bandwidth usage; this valuedepends on the relative frequencies and sizes of docu-ments and queries, and the number ofnodes.We briefly review our assumptions for the analysis.

188宝金博页面版:关注我们

  • 新浪微博

关注188宝金博页面版公众号

188宝金博页面版
阅读
APP
阅读
返回
顶部
188宝金博页面版官网登录在线平台入口(2026已更新)—江苏协昌电子科技股份有限公司