返回信息流From Science Mag.
Science 15 August 2008:
Vol. 321. no. 5891, p. 912
DOI: 10.1126/science.1157523
Prev | Table of Contents | Next
Technical Comments
Comment on "100% Accuracy in Automatic Face Recognition"
Weihong Deng,* Jun Guo, Jiani Hu, Honggang Zhang
Jenkins and Burton (Brevia, 25 January 2008, p. 435) reported that image averaging increased the accuracy of the automatic face recognition to 100% and thus could be applied to photo-identification documents. We argue that the feasibility of image averaging on identification documents is not fully supported by the presented evidence.
School of Information Engineering, Beijing University of Posts and Telecommunications, Beijing, 100876, China.
* To whom correspondence should be addressed. E-mail: whdeng@bupt.edu.cn
In automatic face recognition, a gallery of facial images is first enrolled and coded for subsequent searching. A probe image is then obtained and compared with each encoded face in the gallery, and a recognition is noted when a suitable match occurs. In a recent study, Jenkins and Burton (1) used the photographs of celebrities as probe images to measure the hit rate of the FaceVACS (Cognitec Systems GmbH, Dresden, Germany) face-recognition system used by the genealogy Web site MyHeritage (2). Merging the probe images to create an average image for each celebrity raised the overall hit rate for the probe database from 54 to 100%. The authors therefore concluded that the process of image averaging could dramatically boost automatic face recognition and inferred that incorporating average face images into identification documents would greatly reduce the incidence of face-recognition errors.
As Jenkins and Burton suggest in (1), it is possible that 100% accuracy was achieved simply because the image averages incorporated some recognizable photos. To allay that concern, the authors reported that a new set of averages using only those photographs that were unrecognized in their first study raised the hit rate from 0 to 80%. They thus reasoned that the improved accuracy could solely be attributed to the averaging process. However, the improvement on the hit rate could be partly attributed to the manual facial registration [see supporting online material for (1)] before averaging, which accurately rectified the facial appearance so that all the probe faces were aligned in a standard frontal and upright posture and enclosed by a uniform background. By largely reducing the image variability, the image registration procedure might transform some unrecognized photos into recognizable faces (3). Moreover, the standard registered faces might facilitate the automatic face finding and normalization process of the tested algorithm, which may have also boosted the hit rate (4). It is thus possible that the registration technique assisted the image averaging to boost the hit rate to a higher level.
Jenkins and Burton correctly suggested that image averaging enhanced the performance by stabilizing the face image. However, the interpretation of this fact was overextended. The conclusion that including average images on identification documents would reduce recognition errors lacks sufficient evidence, especially because it is not an equivalent task to the experiments that were carried out. Specifically, the experiments in (1) used the online database as the gallery and the average images as the probe, and the online recognition system only returned the closest matching photo from its database. If the identity of the returned photo matched that of the average image, it was recorded as a hit. Using this methodology, even the 100% hit rate could only ensure that, for each test identity, the system successfully matched the average with "one" gallery photo from that person. However, there were multiple (from 7 to 28) gallery photos for each test identity in the database (1). The experiments did not show the number of single (gallery) photos to which the averages could be matched. In contrast, if the average image is incorporated in identification documents, the identity-verification system must be required to suitably match it to every photo from the same person; otherwise any miss on a photo would translate to a recognition error. Therefore, although the recognition algorithm is commutable, the task of identity verification is more demanding than that of Jenkins and Burton's experiments, and the feasibility of using average images for verifying identity requires further testing. Proof of identity is achieved by comparing an individual's appearance to a photo-identification document, where the appearance is captured by any single facial image in diverse locales and different times. The reliability of the proof depends on how stably the single images can be matched to the corresponding photo-identification document. Hence, in order to evaluate the feasibility of average images on identification documents, a refined experiment should be designed to measure the hit rate for single (gallery) photos, showing what proportion of the single images can be matched to the corresponding average. Moreover, the reliability of the proof also depends on the ability to reject the photos of the impostors according to the averages, which also need to be considered. For the scientific methodology, one can refer to United States government–sponsored evaluations, such as the Face Recognition Vendor Test (5), which are the standard test beds for face-recognition technologies.
We acknowledge that image averaging contributes to the face-recognition procedures. However, automatic face recognition is a complex pattern-recognition problem involved with early processing, perceptual coding, and cue-fusion mechanisms (6). Although countless solid contributions have been made (7), 100% accuracy in automatic face recognition in real-world settings remains an ambitious goal.
References and Notes
1. R. Jenkins, A. M. Burton, Science 319, 435 (2008).[Abstract/Free Full Text]
2. MyHeritage, www.myheritage.com/face-recognition.
3. I. Craw, N. Costen, T. Kato, S. Akamatsu, IEEE Trans. Pattern Anal. 21, 725 (1999). [CrossRef]
4. S. Shan, Y. Chang, W. Gao, B. Cao, Proc. 6th IEEE Int. Conf. Automatic Face and Gesture Recognition, 314 (Seoul, Korea, 17 to 19 May 2004).
5. Face Recognition Vendor Test, www.frvt.org.
6. P. Sinha, Nat. Neurosci. 5, 1093 (2002). [CrossRef] [ISI] [Medline]
7. W. Zhao, R. Chellappa, P. J. Phillips, A. Rosenfield, Assoc. Comput. Mach. Comput. Surv. 35, 399 (2003).
8. The authors are funded by China Scholarship Council, Natural Science Foundation of China (60675001), and National High-Tech Development Plan of China (2007AA01Z417).
Received for publication 10 March 2008. Accepted for publication 15 July 2008.
这是一条镜像帖。来源:北邮人论坛 / ml-dm / #3026同步于 2008/8/30
该镜像源已超过 30 天没有更新,可能在源站已被删除。
ML_DM机器人发帖
fighting on face recognition
cryppie
2008/8/30镜像同步11 回复
订阅后,新回复会通过你的通知中心匿名送达。
9 条回复
response from original authors
Response to Comment on "100% Accuracy in Automatic Face Recognition"
R. Jenkins* and A. M. Burton
Contrary to the suggestion of Deng et al., image registration reduced face-recognition accuracy when divorced from the averaging procedure. Average-to-photo mapping generalizes beyond specific photographs, and averaging either gallery images or probe images can improve the match. The alternative protocol suggested by the authors is unsuitable because it evaluates face-matching algorithms, not face representations, and relies on standard image sets.
Department of Psychology, University of Glasgow, Glasgow G12 8QQ, UK.
* To whom correspondence should be addressed. E-mail: rob@psy.gla.ac.uk
We reported that the process of image averaging can dramatically boost automatic face recognition (1). Deng et al. (2) suggest that image registration alone might improve face-recognition performance, and we tested this suggestion. Because the MyHeritage database (3) is constantly expanding, we first re-submitted the photographs and average images used in (1) to establish a current baseline. Forty-eight of the 500 probe images were identical to images in the online gallery, compared with 41 in (1). This increase is consistent with gallery expansion. Of the remaining 452 photographs, 52% were correctly identified, down from 54% in (1). The hit rate for the average images was 100%, as before. Five of the average images matched different photos of the correct person, confirming that the average-to-photo mapping generalizes beyond particular snapshots. To address Deng et al.'s concern, we next submitted manually registered versions of the source photographs. As Deng et al. describe, these were aligned in a standard frontal and upright posture and enclosed by a uniform background. The hit rate for the registered images was 30%. Apparently, registration alone offers the worst of both worlds: It disrupts any informative correspondence in shape between gallery and probe items but does not otherwise stabilize image variability. Registration of the probe images might be less harmful when the gallery images are also registered. In a previous study using a principal components analysis–based image match (4), we carried out exactly this transformation. Performance was poor but was nonetheless improved by averaging.
Deng et al. (2) also express concern that our average images were presented as probes rather than being gallery items. This was a consequence of our chosen methodology. To ensure a stringent test of our averaging technique, we relinquished control over several key aspects of the image match. We used someone else's gallery photographs together with someone else's matching algorithm. Our probe images were collected from the Internet. This approach meant that we were not able to add images to the gallery, but we could still submit images as probes. Because face recognition can be reduced to matching pairs of images, the order of each pair was not our main interest, and we treated matching A to B as equivalent to matching B to A. In previous studies, we have shown that averaging also helps when applied to the gallery images (4). Whether identity checks would be better served by an average image stored in an identification document or an average probe generated from the live face is an interesting empirical question. However, it is worth pointing out that averaging probe images specifically finds practical application in forensic face recognition (5).
Deng et al. point out that an average probe need only match one gallery photo of the target to score a hit. The same is true for the photographic probes, yet these performed comparatively poorly. In practice, an average probe can match very different photos of the target, as our new data confirm. This underscores the major benefit of averaging. Matching pairs of photos is extremely difficult, because both items contain information that is not diagnostic of identity. Matching a photo to an average is helpful because it eliminates non-diagnostic information from one item in the pair. There is no doubt that difficulties can still arise in this situation, but this is partly because the pair still includes a photograph. Our response is therefore not to retreat to matching pairs of photos but rather to investigate ways to eliminate photos from the match altogether. Matching pairs of average images is an obvious route to explore, and we are testing this possibility.
Deng et al. recommend the Face Recognition Vendor Test (FRVT) (6) as a methodological template. This is unsuitable for several reasons. First, the FRVT evaluations compare performance of different matching algorithms on standard images. Our proposal concerns the representation of the face and is independent of the matching algorithm. Second, the standard databases consist of posed photographs, which grossly underrepresent the variability of ambient face images. Third, reliance on any standard database carries the risk of solving "database recognition" without tackling face recognition. The real world presents different crowds on different days, and systems aspiring to real-world application cannot ignore this inconvenience.
Finally, we agree with Deng et al. that early processing and automatic feature extraction are interesting problems, but they are clearly separate from the problem of face recognition. To convince yourself of this, note that it is easy to locate landmarks on a face you cannot recognize and that doing so does not trigger identification.
References and Notes
1. R. Jenkins, A. M. Burton, Science 319, 435 (2008).[Abstract/Free Full Text]
2. W. Deng, J. Guo, J. Hu, H. Zhang, Science 321, 912 (2008); www.sciencemag.org/cgi/content/full/321/5891/912c.
3. MyHeritage, www.myheritage.com.
4. A. M. Burton, R. Jenkins, P. J. B. Hancock, D. White, Cognit. Psychol. 51, 256 (2005). [CrossRef] [ISI] [Medline]
5. V. Bruce, H. Ness, P. J. B. Hancock, C. Newman, J. Rarity, J. Appl. Psychol. 87, 894 (2002). [CrossRef] [ISI] [Medline]
6. Face Recognition Vendor Test, www.frvt.org.
Received for publication 21 April 2008. Accepted for publication 16 July 2008.
热情一顶
【 在 cryppie (E.Coli) 的大作中提到: 】
: From Science Mag.
: Science 15 August 2008:
: Vol. 321. no. 5891, p. 912
: ...................