+ All Categories
Home > Documents > Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Date post: 05-Aug-2020
Category:
Upload: others
View: 2 times
Download: 0 times
Share this document with a friend
74
Informatica MDM Multidomain Edition (Version 10.1.0) Cleanse Adapter Guide
Transcript
Page 1: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Informatica MDM Multidomain Edition(Version 10.1.0)

Cleanse Adapter Guide

Page 2: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Informatica MDM Multidomain Edition Cleanse Adapter Guide

Version 10.1.0November 2015

Copyright (c) 1993-2015 Informatica LLC. All rights reserved.

This software and documentation contain proprietary information of Informatica LLC and are provided under a license agreement containing restrictions on use and disclosure and are also protected by copyright law. Reverse engineering of the software is prohibited. No part of this document may be reproduced or transmitted in any form, by any means (electronic, photocopying, recording or otherwise) without prior consent of Informatica LLC. This Software may be protected by U.S. and/or international Patents and other Patents Pending.

Use, duplication, or disclosure of the Software by the U.S. Government is subject to the restrictions set forth in the applicable software license agreement and as provided in DFARS 227.7202-1(a) and 227.7702-3(a) (1995), DFARS 252.227-7013©(1)(ii) (OCT 1988), FAR 12.212(a) (1995), FAR 52.227-19, or FAR 52.227-14 (ALT III), as applicable.

The information in this product or documentation is subject to change without notice. If you find any problems in this product or documentation, please report them to us in writing.

Informatica, Informatica Platform, Informatica Data Services, PowerCenter, PowerCenterRT, PowerCenter Connect, PowerCenter Data Analyzer, PowerExchange, PowerMart, Metadata Manager, Informatica Data Quality, Informatica Data Explorer, Informatica B2B Data Transformation, Informatica B2B Data Exchange Informatica On Demand, Informatica Identity Resolution, Informatica Application Information Lifecycle Management, Informatica Complex Event Processing, Ultra Messaging and Informatica Master Data Management are trademarks or registered trademarks of Informatica LLC in the United States and in jurisdictions throughout the world. All other company and product names may be trade names or trademarks of their respective owners.

Portions of this software and/or documentation are subject to copyright held by third parties, including without limitation: Copyright DataDirect Technologies. All rights reserved. Copyright © Sun Microsystems. All rights reserved. Copyright © RSA Security Inc. All Rights Reserved. Copyright © Ordinal Technology Corp. All rights reserved.Copyright © Aandacht c.v. All rights reserved. Copyright Genivia, Inc. All rights reserved. Copyright Isomorphic Software. All rights reserved. Copyright © Meta Integration Technology, Inc. All rights reserved. Copyright © Intalio. All rights reserved. Copyright © Oracle. All rights reserved. Copyright © Adobe Systems Incorporated. All rights reserved. Copyright © DataArt, Inc. All rights reserved. Copyright © ComponentSource. All rights reserved. Copyright © Microsoft Corporation. All rights reserved. Copyright © Rogue Wave Software, Inc. All rights reserved. Copyright © Teradata Corporation. All rights reserved. Copyright © Yahoo! Inc. All rights reserved. Copyright © Glyph & Cog, LLC. All rights reserved. Copyright © Thinkmap, Inc. All rights reserved. Copyright © Clearpace Software Limited. All rights reserved. Copyright © Information Builders, Inc. All rights reserved. Copyright © OSS Nokalva, Inc. All rights reserved. Copyright Edifecs, Inc. All rights reserved. Copyright Cleo Communications, Inc. All rights reserved. Copyright © International Organization for Standardization 1986. All rights reserved. Copyright © ej-technologies GmbH. All rights reserved. Copyright © Jaspersoft Corporation. All rights reserved. Copyright © International Business Machines Corporation. All rights reserved. Copyright © yWorks GmbH. All rights reserved. Copyright © Lucent Technologies. All rights reserved. Copyright (c) University of Toronto. All rights reserved. Copyright © Daniel Veillard. All rights reserved. Copyright © Unicode, Inc. Copyright IBM Corp. All rights reserved. Copyright © MicroQuill Software Publishing, Inc. All rights reserved. Copyright © PassMark Software Pty Ltd. All rights reserved. Copyright © LogiXML, Inc. All rights reserved. Copyright © 2003-2010 Lorenzi Davide, All rights reserved. Copyright © Red Hat, Inc. All rights reserved. Copyright © The Board of Trustees of the Leland Stanford Junior University. All rights reserved. Copyright © EMC Corporation. All rights reserved. Copyright © Flexera Software. All rights reserved. Copyright © Jinfonet Software. All rights reserved. Copyright © Apple Inc. All rights reserved. Copyright © Telerik Inc. All rights reserved. Copyright © BEA Systems. All rights reserved. Copyright © PDFlib GmbH. All rights reserved. Copyright © Orientation in Objects GmbH. All rights reserved. Copyright © Tanuki Software, Ltd. All rights reserved. Copyright © Ricebridge. All rights reserved. Copyright © Sencha, Inc. All rights reserved. Copyright © Scalable Systems, Inc. All rights reserved. Copyright © jQWidgets. All rights reserved. Copyright © Tableau Software, Inc. All rights reserved. Copyright© MaxMind, Inc. All Rights Reserved. Copyright © TMate Software s.r.o. All rights reserved. Copyright © MapR Technologies Inc. All rights reserved. Copyright © Amazon Corporate LLC. All rights reserved. Copyright © Highsoft. All rights reserved. Copyright © Python Software Foundation. All rights reserved. Copyright © BeOpen.com. All rights reserved. Copyright © CNRI. All rights reserved.

This product includes software developed by the Apache Software Foundation (http://www.apache.org/), and/or other software which is licensed under various versions of the Apache License (the "License"). You may obtain a copy of these Licenses at http://www.apache.org/licenses/. Unless required by applicable law or agreed to in writing, software distributed under these Licenses is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the Licenses for the specific language governing permissions and limitations under the Licenses.

This product includes software which was developed by Mozilla (http://www.mozilla.org/), software copyright The JBoss Group, LLC, all rights reserved; software copyright © 1999-2006 by Bruno Lowagie and Paulo Soares and other software which is licensed under various versions of the GNU Lesser General Public License Agreement, which may be found at http:// www.gnu.org/licenses/lgpl.html. The materials are provided free of charge by Informatica, "as-is", without warranty of any kind, either express or implied, including but not limited to the implied warranties of merchantability and fitness for a particular purpose.

The product includes ACE(TM) and TAO(TM) software copyrighted by Douglas C. Schmidt and his research group at Washington University, University of California, Irvine, and Vanderbilt University, Copyright (©) 1993-2006, all rights reserved.

This product includes software developed by the OpenSSL Project for use in the OpenSSL Toolkit (copyright The OpenSSL Project. All Rights Reserved) and redistribution of this software is subject to terms available at http://www.openssl.org and http://www.openssl.org/source/license.html.

This product includes Curl software which is Copyright 1996-2013, Daniel Stenberg, <[email protected]>. All Rights Reserved. Permissions and limitations regarding this software are subject to terms available at http://curl.haxx.se/docs/copyright.html. Permission to use, copy, modify, and distribute this software for any purpose with or without fee is hereby granted, provided that the above copyright notice and this permission notice appear in all copies.

The product includes software copyright 2001-2005 (©) MetaStuff, Ltd. All Rights Reserved. Permissions and limitations regarding this software are subject to terms available at http://www.dom4j.org/ license.html.

The product includes software copyright © 2004-2007, The Dojo Foundation. All Rights Reserved. Permissions and limitations regarding this software are subject to terms available at http://dojotoolkit.org/license.

This product includes ICU software which is copyright International Business Machines Corporation and others. All rights reserved. Permissions and limitations regarding this software are subject to terms available at http://source.icu-project.org/repos/icu/icu/trunk/license.html.

This product includes software copyright © 1996-2006 Per Bothner. All rights reserved. Your right to use such materials is set forth in the license which may be found at http:// www.gnu.org/software/ kawa/Software-License.html.

This product includes OSSP UUID software which is Copyright © 2002 Ralf S. Engelschall, Copyright © 2002 The OSSP Project Copyright © 2002 Cable & Wireless Deutschland. Permissions and limitations regarding this software are subject to terms available at http://www.opensource.org/licenses/mit-license.php.

This product includes software developed by Boost (http://www.boost.org/) or under the Boost software license. Permissions and limitations regarding this software are subject to terms available at http:/ /www.boost.org/LICENSE_1_0.txt.

This product includes software copyright © 1997-2007 University of Cambridge. Permissions and limitations regarding this software are subject to terms available at http:// www.pcre.org/license.txt.

This product includes software copyright © 2007 The Eclipse Foundation. All Rights Reserved. Permissions and limitations regarding this software are subject to terms available at http:// www.eclipse.org/org/documents/epl-v10.php and at http://www.eclipse.org/org/documents/edl-v10.php.

Page 3: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

This product includes software licensed under the terms at http://www.tcl.tk/software/tcltk/license.html, http://www.bosrup.com/web/overlib/?License, http://www.stlport.org/doc/ license.html, http://asm.ow2.org/license.html, http://www.cryptix.org/LICENSE.TXT, http://hsqldb.org/web/hsqlLicense.html, http://httpunit.sourceforge.net/doc/ license.html, http://jung.sourceforge.net/license.txt , http://www.gzip.org/zlib/zlib_license.html, http://www.openldap.org/software/release/license.html, http://www.libssh2.org, http://slf4j.org/license.html, http://www.sente.ch/software/OpenSourceLicense.html, http://fusesource.com/downloads/license-agreements/fuse-message-broker-v-5-3- license-agreement; http://antlr.org/license.html; http://aopalliance.sourceforge.net/; http://www.bouncycastle.org/licence.html; http://www.jgraph.com/jgraphdownload.html; http://www.jcraft.com/jsch/LICENSE.txt; http://jotm.objectweb.org/bsd_license.html; . http://www.w3.org/Consortium/Legal/2002/copyright-software-20021231; http://www.slf4j.org/license.html; http://nanoxml.sourceforge.net/orig/copyright.html; http://www.json.org/license.html; http://forge.ow2.org/projects/javaservice/, http://www.postgresql.org/about/licence.html, http://www.sqlite.org/copyright.html, http://www.tcl.tk/software/tcltk/license.html, http://www.jaxen.org/faq.html, http://www.jdom.org/docs/faq.html, http://www.slf4j.org/license.html; http://www.iodbc.org/dataspace/iodbc/wiki/iODBC/License; http://www.keplerproject.org/md5/license.html; http://www.toedter.com/en/jcalendar/license.html; http://www.edankert.com/bounce/index.html; http://www.net-snmp.org/about/license.html; http://www.openmdx.org/#FAQ; http://www.php.net/license/3_01.txt; http://srp.stanford.edu/license.txt; http://www.schneier.com/blowfish.html; http://www.jmock.org/license.html; http://xsom.java.net; http://benalman.com/about/license/; https://github.com/CreateJS/EaselJS/blob/master/src/easeljs/display/Bitmap.js; http://www.h2database.com/html/license.html#summary; http://jsoncpp.sourceforge.net/LICENSE; http://jdbc.postgresql.org/license.html; http://protobuf.googlecode.com/svn/trunk/src/google/protobuf/descriptor.proto; https://github.com/rantav/hector/blob/master/LICENSE; http://web.mit.edu/Kerberos/krb5-current/doc/mitK5license.html; http://jibx.sourceforge.net/jibx-license.html; https://github.com/lyokato/libgeohash/blob/master/LICENSE; https://github.com/hjiang/jsonxx/blob/master/LICENSE; https://code.google.com/p/lz4/; https://github.com/jedisct1/libsodium/blob/master/LICENSE; http://one-jar.sourceforge.net/index.php?page=documents&file=license; https://github.com/EsotericSoftware/kryo/blob/master/license.txt; http://www.scala-lang.org/license.html; https://github.com/tinkerpop/blueprints/blob/master/LICENSE.txt; http://gee.cs.oswego.edu/dl/classes/EDU/oswego/cs/dl/util/concurrent/intro.html; https://aws.amazon.com/asl/; https://github.com/twbs/bootstrap/blob/master/LICENSE; https://sourceforge.net/p/xmlunit/code/HEAD/tree/trunk/LICENSE.txt; https://github.com/documentcloud/underscore-contrib/blob/master/LICENSE, and https://github.com/apache/hbase/blob/master/LICENSE.txt.

This product includes software licensed under the Academic Free License (http://www.opensource.org/licenses/afl-3.0.php), the Common Development and Distribution License (http://www.opensource.org/licenses/cddl1.php) the Common Public License (http://www.opensource.org/licenses/cpl1.0.php), the Sun Binary Code License Agreement Supplemental License Terms, the BSD License (http:// www.opensource.org/licenses/bsd-license.php), the new BSD License (http://opensource.org/licenses/BSD-3-Clause), the MIT License (http://www.opensource.org/licenses/mit-license.php), the Artistic License (http://www.opensource.org/licenses/artistic-license-1.0) and the Initial Developer’s Public License Version 1.0 (http://www.firebirdsql.org/en/initial-developer-s-public-license-version-1-0/).

This product includes software copyright © 2003-2006 Joe WaInes, 2006-2007 XStream Committers. All rights reserved. Permissions and limitations regarding this software are subject to terms available at http://xstream.codehaus.org/license.html. This product includes software developed by the Indiana University Extreme! Lab. For further information please visit http://www.extreme.indiana.edu/.

This product includes software Copyright (c) 2013 Frank Balluffi and Markus Moeller. All rights reserved. Permissions and limitations regarding this software are subject to terms of the MIT license.

See patents at https://www.informatica.com/legal/patents.html.

DISCLAIMER: Informatica LLC provides this documentation "as is" without warranty of any kind, either express or implied, including, but not limited to, the implied warranties of noninfringement, merchantability, or use for a particular purpose. Informatica LLC does not warrant that this software or documentation is error free. The information provided in this software or documentation may include technical inaccuracies or typographical errors. The information in this software and documentation is subject to change at any time without notice.

NOTICES

This Informatica product (the "Software") includes certain drivers (the "DataDirect Drivers") from DataDirect Technologies, an operating company of Progress Software Corporation ("DataDirect") which are subject to the following terms and conditions:

1.THE DATADIRECT DRIVERS ARE PROVIDED "AS IS" WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESSED OR IMPLIED, INCLUDING BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NON-INFRINGEMENT.

2. IN NO EVENT WILL DATADIRECT OR ITS THIRD PARTY SUPPLIERS BE LIABLE TO THE END-USER CUSTOMER FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, CONSEQUENTIAL OR OTHER DAMAGES ARISING OUT OF THE USE OF THE ODBC DRIVERS, WHETHER OR NOT INFORMED OF THE POSSIBILITIES OF DAMAGES IN ADVANCE. THESE LIMITATIONS APPLY TO ALL CAUSES OF ACTION, INCLUDING, WITHOUT LIMITATION, BREACH OF CONTRACT, BREACH OF WARRANTY, NEGLIGENCE, STRICT LIABILITY, MISREPRESENTATION AND OTHER TORTS.

Part Number: MDM-CAG-10100-0001

Page 4: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Table of Contents

Preface . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7Informatica Resources. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

Informatica My Support Portal. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

Informatica Documentation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

Informatica Product Availability Matrixes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

Informatica Web Site. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

Informatica How-To Library. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

Informatica Knowledge Base. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

Informatica Support YouTube Channel. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

Informatica Marketplace. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

Informatica Velocity. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8

Informatica Global Customer Support. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9

Chapter 1: Introduction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10Supported Cleanse Engines. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10

Prerequisites. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10

Chapter 2: Informatica IDQ Cleanse Engine. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12Informatica IDQ Cleanse Engine Overview. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12

Prerequisites for Data Cleansing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12

Process Overview. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13

Run-time Behavior. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13

Considerations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14

Adding an IDQ Library in the Cleanse Functions Tool. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15

Configuring Generated Libraries. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17

Enabling a Generated Library. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17

Properties File Syntax. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17

Example Properties File. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18

Editing a Generated Properties File. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18

Updating the Connection Endpoints for Web Services. . . . . . . . . . . . . . . . . . . . . . . . . . . 21

Chapter 3: AddressDoctor Cleanse Engine. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23AddressDoctor Cleanse Engine Overview. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23

Configuring the AddressDoctor Cleanse Engine. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23

Obtaining the Required Files. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24

Editing the Properties File. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24

Upgrading the AddressDoctor Cleanse Engine. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26

Obtaining the Required Files. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26

Editing the Properties File. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26

Replacing AddressDoctor 4 Library with AddressDoctor 5 Library. . . . . . . . . . . . . . . . . . . . 26

4 Table of Contents

Page 5: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Steps to Upgrade to AddressDoctor 5 Integration. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27

AddressDoctor 5 Fields and Process Status Values. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29

AddressDoctor 5 Input Fields. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29

AddressDoctor 5 Output Fields. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32

Address Examples. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39

Differences Between AddressDoctor 4 and AddressDoctor 5 Process Status Values. . . . . . . 43

Configuring the JVM Settings. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45

Setting the JVM Size for WebSphere on Windows/UNIX. . . . . . . . . . . . . . . . . . . . . . . . . . 45

Chapter 4: FirstLogic Direct Cleanse Engine. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47FirstLogic Direct Cleanse Engine Overview. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47

About FirstLogic Direct Integration. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47

Installing the Components. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48

Configuring FirstLogic Direct. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48

Configuring Informatica MDM Hub to Use the Adapter. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48

Using Your FirstLogic Direct Library. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49

About Transactional Mode. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49

About the Transactional Mode Sample. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49

Chapter 5: Trillium Director Cleanse Engine. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51Trillium Director Cleanse Engine Overview. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51

About Trillium Director Integration. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51

Before You Install. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52

Configuring Trillium Director and the Cleanse Match Server. . . . . . . . . . . . . . . . . . . . . . . . . . 52

Testing Your Trillium Director Configuration. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54

Sample Configuration Files. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54

Upgrading Trillium Director. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56

Using Trillium on a Remote Server. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56

Configuring Trillium Director for Multithreading. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57

Setting the Threading Pool. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58

Increasing the Number of Network Connection Retries. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58

Chapter 6: SAP Data Services XI Cleanse Engine. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59SAP Data Services XI Cleanse Engine Overview. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59

About SAP Data Services XI Integration. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59

Prerequisites. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59

Process Overview. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60

Run-time Behavior. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61

Considerations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61

Adding an SAP Library in the Cleanse Functions Tool. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61

Configuring Generated Libraries. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64

Enabling a Generated Library. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64

Properties File Syntax. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64

Table of Contents 5

Page 6: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Example Properties File. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64

Editing a Generated Properties File. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65

Chapter 7: Troubleshooting. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69AddressDoctor Initialization Issues. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69

Cleanse Engine Initialization Fails. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69

Remote Initialization Fails. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70

Trillium Errors. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70

Initialization Fails for MDM Console. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70

Remote Initialization Fails. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71

Trillium Director Integration. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71

About Trillium Director Working Files. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71

Finding the Location of the Trillium Director Work Files. . . . . . . . . . . . . . . . . . . . . . . . . . 71

Setting the Location of the Process Server Work Files. . . . . . . . . . . . . . . . . . . . . . . . . . . 71

Setting Whether Working Files are Kept. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71

Setting the Number of Connections in a Connection Pool. . . . . . . . . . . . . . . . . . . . . . . . . 72

Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73

6 Table of Contents

Page 7: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

PrefaceWelcome to the Informatica MDM Hub Cleanse Adapter Guide. This guide describes how to configure cleanse engines to work with the Informatica MDM Hub and Informatica MDM adapters. The guide also provides prerequisites and test procedures for each adapter.

This guide has been written for database administrators, system administrators, and other implementers who are responsible for the setup tasks required for cleanse adapters and engines. System administrators must be familiar with the Windows and UNIX platforms. Knowledge of Oracle administration is particularly important.

Other administration and configuration tasks are described in the Informatica MDM Hub Configuration Guide .

Informatica Resources

Informatica My Support PortalAs an Informatica customer, the first step in reaching out to Informatica is through the Informatica My Support Portal at https://mysupport.informatica.com. The My Support Portal is the largest online data integration collaboration platform with over 100,000 Informatica customers and partners worldwide.

As a member, you can:

• Access all of your Informatica resources in one place.

• Review your support cases.

• Search the Knowledge Base, find product documentation, access how-to documents, and watch support videos.

• Find your local Informatica User Group Network and collaborate with your peers.

Informatica DocumentationThe Informatica Documentation team makes every effort to create accurate, usable documentation. If you have questions, comments, or ideas about this documentation, contact the Informatica Documentation team through email at [email protected]. We will use your feedback to improve our documentation. Let us know if we can contact you regarding your comments.

The Documentation team updates documentation as needed. To get the latest documentation for your product, navigate to Product Documentation from https://mysupport.informatica.com.

7

Page 8: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Informatica Product Availability MatrixesProduct Availability Matrixes (PAMs) indicate the versions of operating systems, databases, and other types of data sources and targets that a product release supports. You can access the PAMs on the Informatica My Support Portal at https://mysupport.informatica.com.

Informatica Web SiteYou can access the Informatica corporate web site at https://www.informatica.com. The site contains information about Informatica, its background, upcoming events, and sales offices. You will also find product and partner information. The services area of the site includes important information about technical support, training and education, and implementation services.

Informatica How-To LibraryAs an Informatica customer, you can access the Informatica How-To Library at https://mysupport.informatica.com. The How-To Library is a collection of resources to help you learn more about Informatica products and features. It includes articles and interactive demonstrations that provide solutions to common problems, compare features and behaviors, and guide you through performing specific real-world tasks.

Informatica Knowledge BaseAs an Informatica customer, you can access the Informatica Knowledge Base at https://mysupport.informatica.com. Use the Knowledge Base to search for documented solutions to known technical issues about Informatica products. You can also find answers to frequently asked questions, technical white papers, and technical tips. If you have questions, comments, or ideas about the Knowledge Base, contact the Informatica Knowledge Base team through email at [email protected].

Informatica Support YouTube ChannelYou can access the Informatica Support YouTube channel at http://www.youtube.com/user/INFASupport. The Informatica Support YouTube channel includes videos about solutions that guide you through performing specific tasks. If you have questions, comments, or ideas about the Informatica Support YouTube channel, contact the Support YouTube team through email at [email protected] or send a tweet to @INFASupport.

Informatica MarketplaceThe Informatica Marketplace is a forum where developers and partners can share solutions that augment, extend, or enhance data integration implementations. By leveraging any of the hundreds of solutions available on the Marketplace, you can improve your productivity and speed up time to implementation on your projects. You can access Informatica Marketplace at http://www.informaticamarketplace.com.

Informatica VelocityYou can access Informatica Velocity at https://mysupport.informatica.com. Developed from the real-world experience of hundreds of data management projects, Informatica Velocity represents the collective knowledge of our consultants who have worked with organizations from around the world to plan, develop, deploy, and maintain successful data management solutions. If you have questions, comments, or ideas about Informatica Velocity, contact Informatica Professional Services at [email protected].

8 Preface

Page 9: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Informatica Global Customer SupportYou can contact a Customer Support Center by telephone or through the Online Support.

Online Support requires a user name and password. You can request a user name and password at http://mysupport.informatica.com.

The telephone numbers for Informatica Global Customer Support are available from the Informatica web site at http://www.informatica.com/us/services-and-training/support-services/global-support-centers/.

Preface 9

Page 10: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

C H A P T E R 1

IntroductionThis chapter includes the following topics:

• Supported Cleanse Engines, 10

• Prerequisites, 10

Supported Cleanse EnginesThe following table displays the cleanse engines that Informatica MDM Hub supports and the Informatica MDM adapters that they work with:

Cleanse Engine Informatica MDM Hub Adapter

IDQ Informatica IDQ Adapter

AddressDoctor AddressDoctor Adapter

FirstLogic Direct FirstLogic Data Quality Adapter

Trillium Trillium Director Adapter

SAP Data Services XI SAP Data Services XI Adapter

PrerequisitesBefore you can use your cleanse engines, you might need to do the following, depending on which cleanse engine you are using:

• Obtain a license with cleanse adapter enabled from Informatica.

• Install the Hub, application server, Process Server, and cleanse adapter. In some cases, you must install the Process Server and the cleanse adapter on the same server.

• Configure the cleanse engine and Informatica MDM Hub. See individual chapters of this guide more information.

• Test your cleanse-engine configuration.

For more information related to your cleanse adapter configuration, see your third party cleanse engine documentation, your application server documentation, the Informatica MDM Hub Installation Guide, and the

10

Page 11: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Informatica MDM Hub release notes. You can also check the Informatica Knowledge Base for information about specific issues.

Prerequisites 11

Page 12: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

C H A P T E R 2

Informatica IDQ Cleanse EngineThis chapter includes the following topics:

• Informatica IDQ Cleanse Engine Overview, 12

• Adding an IDQ Library in the Cleanse Functions Tool, 15

• Configuring Generated Libraries, 17

Informatica IDQ Cleanse Engine OverviewOne of the ways to access data cleansing functionality in the Informatica Data Quality product is through Web services that Informatica publishes. This chapter describes how, in your Informatica MDM Hub implementation, to set up and use the MDM Cleanse Adapter for Informatica Data Quality (Web Services) to access cleanse functions that are published as Web services.

This functionality allows you add a new type of cleanse library - an IDQ cleanse library - to your Informatica MDM Hub implementation, and then integrate cleanse functions in the IDQ library into your mappings, just as you would integrate any other type of cleanse function available in your Informatica MDM Hub implementation. Informatica MDM Hub acts as a Web service client application that consumes Informatica Web services.

If you do not use the Informatica IDQ cleanse engine, you can skip this chapter.

Prerequisites for Data CleansingTo use this functionality, you must have installed the following software:

• Informatica MDM Multidomain Edition

• Informatica PowerCenter

• Informatica MDM license that enables IDQ cleanse library functionality (siperian.informatica_DQ=yes in the siperian.license file)

Any tools that you use must support the correct protocols and Java version. For more information about product requirements and supported platforms, see the Product Availability Matrix on the Informatica My Support Portal: https://mysupport.informatica.com/community/my-support/product-availability-matrices.

12

Page 13: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Process OverviewTo use these Web services in your Informatica MDM Hub implementation, complete the following steps:

1. Create a mapping in IDQ or PowerCenter. See PowerCenter Web Services Provider Guide available at the Informatica MySupport website for information on how to create Web Services Description Language (WSDL) files for web services. Obtain the Informatica WSDL file for the Web services that you want to consume. Use the WSDL you created in Informatica PowerCenter.

2. In the Cleanse Functions tool of the Informatica MDM Hub Console, add an IDQ library and specify the URI of the Web Services Description Language (WSDL) file as well as connection information (service and port) to hosted Web services.

The Cleanse Functions tool builds the IDQ library based on the WSDL, and then displays the list of cleanse functions defined in the WSDL. Each cleanse function represents a separate Web service.

3. Enable a generated library.

4. In the Mappings tool of the Informatica MDM Hub Console, use the available cleanse functions in your mappings as required. You must configure the inputs and outputs just as you would configure any other mappings.

For performance reasons, the IDQ function should be the only one used in the mapping. You can have direct column-to-column mappings in the same mapping as an IDQ function, but you should not use other functions or conditional execution components in the same mapping as an IDQ function.

Related Topics:• “Adding an IDQ Library in the Cleanse Functions Tool” on page 15

• “Enabling a Generated Library” on page 17

• “Run-time Behavior” on page 13

Run-time BehaviorAt run time, the cleanse engine evaluates the number of records to be cleansed. If only one (1) record is sent for cleansing, then the web service is invoked for that single record. If multiple records are to be cleansed,

Informatica IDQ Cleanse Engine Overview 13

Page 14: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

then the Process Server will batch a set of records together and pass that batch to the web service. This reduces latency on the web service call and results in better performance.

The maximum number of records included in a batch is determined by a parameter in the cmxcleanse.properties file:

cmx.server.cleanse.number_of_recs_batch= nnn

The default value for that parameter is 50.

Considerations• Add the following parameter to the cmxcleanse.properties file if any of your IDQ functions do not

support minibatch:cmx.server.cleanse.number_of_recs_batch=1

• Web service invocations are synchronous only. Asynchronous invocations are not supported.

• By default, web service invocations typically operate on a single record at a time. However, the MDM Hub cleanse engine will batch together records for cleansing to improve the performance. It can only do this if all the data transformation logic resides in the IDQ web service. If any transformation logic is defined in the MDM map, then the cleanse engine will not be able to use batching logic and the calls to the web service revert to being single record invocations of the web service

• The IDQ function must contain all transformation logic. This is so that no other functions are needed in the MDM map, thus allowing batches of records to be passed to IDQ for processing. The purpose of using the Web services is strictly to transform data that is passed in the request according to the associated cleanse function. Other types of Web services, such as publish/subscribe services, are not supported.

• If the Web service returns an error, Informatica MDM Hub moves the record to the reject table and saves a description of the problem (including any error information returned from the Web service).

• If the Web service is published on a remote system, the infrastructure must be in place for Informatica MDM Hub to connect to the Web service (such as a network that accesses the Internet).

• When using cleanse functions that are implemented as Web services, run-time performance of Web service invocations depends on some factors that are external to Informatica MDM Hub, such as availability of the Web service, the time required for the Web service to process the request and return the response, and network speed.

• You can run WSDL cleanse function with multi-threading. To enable this, change the thread count on the Informatica MDM Hub Process Server. Ensure that there is a sufficient number of instances of the IDQ Web Service to handle the multiple Informatica MDM cleanse threads; otherwise, records might be rejected due to timeouts.

• WSDL files must comply with the Axis2 Databinding Framework (ADB). Non-compliant WSDL files are not supported.

• When you configure mappings, you must ensure that the inputs and outputs are appropriate for the Web service you are calling. The Mappings tool does not validate your inputs and outputs – this is done by the Web service instead. If you have invalid inputs or outputs, the Web service returns an error response and processed records are moved to the reject table with an explanation of the error.

14 Chapter 2: Informatica IDQ Cleanse Engine

Page 15: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Adding an IDQ Library in the Cleanse Functions ToolOnce you have installed the prerequisite software and obtained an IDQ WSDL file, you use the Cleanse Functions tool in the Informatica MDM Hub Console to add the IDQ library to your Informatica MDM Hub implementation.

1. Launch the Informatica MDM Hub Console, if it is not already running.

2. Start the Cleanse Functions tool. You can right click anywhere in the Cleanse Functions tool to see more options.

3. Obtain a write lock (Write Lock > Acquire Lock).

4. Select the Cleanse Functions (root) node. Right click.

5. Choose Cleanse Functions > Add IDQ Library.

6. In the Add IDQ Library dialog, specify the following settings:

Setting Description

Library Name Name of this IDQ library. You can assign any arbitrary name that helps you classify and organize the collection of IDQ cleanse functions. Consider having IDQ in the name to distinguish this from other cleanse function libraries. This name appears as the folder name in the Cleanse Functions list.

IDQ WSDL URI URI (location) of the IDQ WSDL to implement.

IDQ WSDL Service Service of the IDQ WSDL to implement.

IDQ WSDL Port Port of the IDQ WSDL to implement.

Description Descriptive text for this library that you want displayed in the Cleanse Functions tool.

Note: Simple WSDLs often have only one Service and one Port. You can refer to the IDQ WSDL for the values to specify for these settings.

The following figure shows sample settings for an IDQ WSDL that invokes default name and address cleansing:

You must ensure that the IDQ WSDL response definition does not specify an array by setting the maxOccurs attribute value to " 1" , as shown in the sample:

<xsd:element minOccurs="1" maxOccurs="1" name="Street" type="xsd:string"/>

Adding an IDQ Library in the Cleanse Functions Tool 15

Page 16: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

7. Click OK to add the metadata definition for this new IDQ library to the local ORS repository.

8. Click the Refresh button to generate the IDQ library.

The Cleanse Functions tool retrieves the latest IDQ WSDL, generates the IDQ library, and displays any available cleanse functions in the Cleanse Functions list.

An error message is displayed in the following cases:

• If the Cleanse Functions tool cannot consume the IDQ WSDL file (for example, due to a syntax error), then it displays an error message instead. You must fix the IDQ WSDL file or obtain a valid one.

Note:

Ensure that the IDQ WSDL file does not contain an array in its response definition. You must set the maxOccurs attribute in the WSDL file to '1'.

• If changes in the IDQ WSDL file affect any existing mappings, the Cleanse Functions tool displays an error. You must fix the affected mappings before running the cleanse process.

16 Chapter 2: Informatica IDQ Cleanse Engine

Page 17: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

9. Click a cleanse function to display its properties.

10. Test the function by clicking the Test tab and then clicking the Test button.

11. At this point, you can add these cleanse functions to your mappings in the Mappings tool, as shown in the “Process Overview” earlier in this chapter.

Configuring Generated LibrariesWhen the Cleanse Functions tool generates a library, it creates a set of properties files in the SIP_HOME/cleanse/lib directory with the following naming format:

siperian-cleanse-<servicename>_<functionname>.properties.new

Note: File names must be unique (service name + function name) within this directory.

Enabling a Generated LibraryTo enable a generated library, remove the .new extension from the file name.

Properties File Syntax# endPointOverride =## Inputs for <functionname>#

Configuring Generated Libraries 17

Page 18: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

#<wsdl_in_param1> = name , description#<wsdl_in_param2> = name , description## Outputs for <functionname>##<wsdl_out_param1> = name , description#<wsdl_out_param2> = name , description

Example Properties FileBelow is an example code listing.

#endPointOverride =## Inputs for name_Address#Req_Name_DITypeVarchar1024 = Req_Name_DITypeVarchar1024 , Req_Name_DITypeVarchar1024Req_State_DITypeVarchar1024 = Req_State_DITypeVarchar1024 ,Req_State_DITypeVarchar1024Req_Address_DITypeVarchar1024 = Req_Address_DITypeVarchar1024 ,Req_Address_DITypeVarchar1024Req_City_DITypeVarchar1024 = Req_City_DITypeVarchar1024 , Req_City_DITypeVarchar1024Req_Pcode_DITypeVarchar1024 = Req_Pcode_DITypeVarchar1024 ,Req_Pcode_DITypeVarchar1024## Outputs for name_Address#Res_Address_DITypeVarchar1024 = Res_Address_DITypeVarchar1024 ,Res_Address_DITypeVarchar1024

Res_Name_DITypeVarchar1024 = Res_Name_DITypeVarchar1024 , Res_Name_DITypeVarchar1024Res_State_DITypeVarchar1024 = Res_State_DITypeVarchar1024 ,Res_State_DITypeVarchar1024Res_City_DITypeVarchar1024 = Res_City_DITypeVarchar1024 , Res_City_DITypeVarchar1024Res_Pcode_DITypeVarchar1024 = Res_Pcode_DITypeVarchar1024 ,Res_Pcode_DITypeVarchar1024

Editing a Generated Properties FileYou need edit a generated properties file for the following reasons:

• to change the connection endpoint

• to rename or remove extraneous parameters

To edit a generated properties file:

1. Remove the .new extension from the end of the file name to enable it.

2. Open the file in a text editor.

3. Make the changes you want to the file, then save the file.

4. In the Cleanse Functions tool, select the IDQ library, then click the Refresh button to activate these changes in the properties file.

Note: When you click Refresh to generate a library from a WSDL, the Cleanse Functions tool regenerates the file with the .new extension. Changes made to a file that has been renamed (without a.new extension) are not overwritten when the library is refreshed.

18 Chapter 2: Informatica IDQ Cleanse Engine

Page 19: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Managing Input and Output ParametersThe names generated from the WSDL can sometimes be long and difficult to read, or in an uncommon order, as shown in the following example.

You can edit the generated parameter file and rename parameters to make them easier to read and recognize. In fact, if you encounter an error indicating that one or more names are too long to be stored in the ORS repository (> 100 characters), then you must shorten these names.

You can simplify the library by removing extraneous parameters that are not needed for your Web service invocations. You can also reorder parameters.

To edit parameter names:

• For required parameters (that you want to use in your cleanse functions), uncomment by removing the # character. Any parameter that is not uncommented is removed from the function.

• Edit parameter names and descriptions as required. You cannot have duplicate parameter names in one function (duplicates are ignored).

The library must be refreshed for the changes to the properties file to take effect. In the Cleanse Functions tool, select the library and click Refresh.

For example, the following figure shows the default generated output for an IDQ library.

Configuring Generated Libraries 19

Page 20: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

To rename a parameter, change its name setting as shown in the following example:

<parameter> = name , description

Save changes to the file and then, in the Cleanse Functions tool, click Refresh to update the modified cleanse function settings.

20 Chapter 2: Informatica IDQ Cleanse Engine

Page 21: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Once refreshed, be sure to test the cleanse function to verify its operation.

Updating the Connection Endpoints for Web ServicesIf the communication endpoint for a Web service changes, or if you must point to a different environment, update the endpoint URL.

1. Start the Hub Console.

2. From the Model workbench, click Cleanse Functions.

The Cleanse Functions tool appears.

3. From the Write Lock menu, click Acquire Lock.

The MDM Hub locks the cleanse libraries for editing.

4. Click the Informatica Data Quality library.

The Informatica Data Quality cleanse library properties appear.

5. Update the Informatica Data Quality cleanse library properties for the web service.

To update the properties for the Informatica Data Quality library, see the property values in the Informatica Data Quality WSDL file.

The following table describes the Informatica Data Quality library properties that you need to update:

Property Name Description

IDQ WSDL URI URL of the Informatica Data Quality WSDL that you updated or implemented.

IDQ WSDL Service Service name of the Informatica Data Quality WSDL that you updated or implemented.

IDQ WSDL Port Port of the Informatica Data Quality WSDL that you updated or implemented.

Configuring Generated Libraries 21

Page 22: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

6. Save the changes and click Refresh.

The Cleanse Functions tool retrieves the latest Informatica Data Quality WSDL and generates the Informatica Data Quality library.

22 Chapter 2: Informatica IDQ Cleanse Engine

Page 23: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

C H A P T E R 3

AddressDoctor Cleanse EngineThis chapter includes the following topics:

• AddressDoctor Cleanse Engine Overview, 23

• Configuring the AddressDoctor Cleanse Engine, 23

• Upgrading the AddressDoctor Cleanse Engine, 26

• Steps to Upgrade to AddressDoctor 5 Integration, 27

• AddressDoctor 5 Fields and Process Status Values, 29

• Configuring the JVM Settings, 45

AddressDoctor Cleanse Engine OverviewIntegration between Informatica MDM Hub and the AddressDoctor cleanse engine occurs through the Informatica MDM Hub AddressDoctor Adapter. This adapter is an optional component.

This chapter explains how to configure your Informatica MDM Hub system to use the AddressDoctor Adapter and the AddressDoctor Cleanse Engine. This chapter assumes that you are knowledgeable about configuring and using the AddressDoctor software.

This chapter also discusses how to obtain the required files, how to edit the properties file, and how to configure the application server JVM settings.

Note: The Informatica MDM Hub requires a very specific AddressDoctor build of the AddressDoctor Cleanse Engine. Check the release notes for the most current information about this build.

Configuring the AddressDoctor Cleanse EngineBefore you can use the AddressDoctor Cleanse Engine, you must complete the following tasks:

1. Contact Informatica Support to obtain the following:

• Correct unlock codes that must be set in the SetConfig.xml file.

• License file (with the AddressDoctor cleanse adapter enabled). The license must include the AddressDoctor adapter.

2. Obtain the required reference data files.

3. Verify the settings in the Parameters.xml file.

23

Page 24: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

4. Edit the cleanse properties file.

5. Add the AddressDoctor library to the PATH (LIBPATH or LD_LIBRARY_PATH depending on your platform) environment variable.

Windows: Add <infamdm_install_directory>\hub\cleanse\lib.

UNIX: Add <infamdm_install_directory>/hub/cleanse/lib.

6. Configure the application server JVM settings.

Note: The Informatica MDM Hub requires a very specific AddressDoctor build of the AddressDoctor Cleanse Engine. For the most current information about this build, see the Product Availability Matrix for the release that is available at Informatica's customer support portal.

Related Topics:• “Configuring the JVM Settings” on page 45

Obtaining the Required FilesYou need to obtain AddressDoctor reference data files for your AddressDoctor cleanse engine configuration that have certifications, such as for the Software Evaluation and Recognition Program (SERP) and Address Matching Approval System (AMAS), and copy them to the AddressDoctor installation directory.

For example,

• Windows: C:\AddressDoctor\5

• UNIX: /u1/addressDoctor/5

Editing the Properties FileTo edit the properties file:

1. Open the cmxcleanse.properties file for editing.

This file is located in:

Windows: < infamdm_install_directory > \hub\cleanse\resourcesUNIX: < infamdm_install_directory > /hub/cleanse/resources

2. Ensure that the following AddressDoctor 5 properties are set in the cmxcleanse.properties files:

Windows:cleanse.library.addressDoctor.property.SetConfigFile=C:/infamdm/hub/cleanse/resources/AddressDoctor/5/SetConfig.xmlcleanse.library.addressDoctor.property.ParametersFile=C:/infamdm/hub/cleanse/resources/AddressDoctor/5/Parameters.xmlcleanse.library.addressDoctor.property.DefaultCorrectionType=PARAMETERS_DEFAULT

UNIX:cleanse.library.addressDoctor.property.SetConfigFile=/u1/infamdm/hub/cleanse/resources/AddressDoctor/5/SetConfig.xmlcleanse.library.addressDoctor.property.ParametersFile=/u1/infamdm/hub/cleanse/resources/AddressDoctor/5/Parameters.xmlcleanse.library.addressDoctor.property.DefaultCorrectionType=PARAMETERS_DEFAULT

3. Copy SetConfig.xml and Parameters.xml to the location specified in the cmxcleanse.properties file.

24 Chapter 3: AddressDoctor Cleanse Engine

Page 25: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

The following is a sample SetConfig.xml file:

<?xml version="1.0" encoding="iso-8859-1"?><SetConfig> <General WriteXMLEncoding="UTF-16" WriteXMLBOM="NEVER" MaxMemoryUsageMB="1024" MaxAddressObjectCount="10" MaxThreadCount="1"/>

<UnlockCode>unlock_code</UnlockCode> <DataBase CountryISO3="ALL" Type="BATCH_INTERACTIVE" Path="<address_doctor_path>" PreloadingType="NONE"/></SetConfig>

The following is a sample Parameters.xml file:

<?xml version="1.0" encoding="iso-8859-1"?><!DOCTYPE Parameters SYSTEM 'Parameters.dtd'><Parameters WriteXMLEncoding="UTF-16" WriteXMLBOM="NEVER"> <Process Mode="BATCH" EnrichmentGeoCoding="ON" EnrichmentCASS="ON" EnrichmentSERP="ON" EnrichmentSNA="ON" EnrichmentSupplementaryGB="ON" EnrichmentSupplementaryUS="ON" /> <Input Encoding="UTF-16" FormatType="ALL" FormatWithCountry="ON" FormatDelimiter="PIPE" /> <Result AddressElements="STANDARD" Encoding="UTF-16" CountryType="NAME_EN" FormatDelimiter="PIPE" /></Parameters>

4. Save and close the properties file.

5. Connect to the database as the schema owner and run the following statement to enable AddressDoctor: update C_REPOS_CL_FUNCTION_LIB set DISPLAY_IND = 1 where function_lib_name = 'AddressDoctor'

6. Restart your application server.

7. Verify that the application server started up properly with no errors.

The AddressDoctor library is displayed in the Hub Console.

8. Update your JVM settings. This increases resources available to the JVM.

Note: You must update AddressDoctor settings in the cmxcleanse.properties file:

• If you move the SetConfig.xml and Parameters.xml files to a location other than that specified during installation.

• If you need to use a different CorrectionType. The supported CorrectionTypes are PARAMETERS_DEFAULT, PARSE_ONLY, CORRECT_ONLY, CERTIFY_ONLY, CORRECT_THEN_CERTIFY, and TRY_CERTIFY_THEN_CORRECT.

Related Topics:• “Troubleshooting” on page 69

• “Configuring the JVM Settings” on page 45

Configuring the AddressDoctor Cleanse Engine 25

Page 26: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Upgrading the AddressDoctor Cleanse EngineIf you are upgrading your Informatica MDM installation, you must also upgrade your AddressDoctor Cleanse Engine. The following sections cover configuration information for upgrading the AddressDoctor Cleanse Engine.

The AddressDoctor Cleanse Engine libraries are included as part of the Informatica MDM Hub upgrade. If you intend to use AddressDoctor Cleanse Engine as your cleanse engine, contact Informatica Global Customer Support for the correct unlock codes.

To configure AddressDoctor:

1. Obtain the required files.

2. Obtain unlock codes that must be set in the SetConfig.xml file.

3. Verify the settings in the Parameters.xml file.

4. Edit the Hub Cleanse Match properties file.

5. Replace the old AddressDoctor library with the new AddressDoctor library.

6. Configure the application server JVM settings.

Related Topics:• “Editing the Properties File” on page 24

• “Configuring the JVM Settings” on page 45

Obtaining the Required FilesYou need to obtain AddressDoctor reference data files for your AddressDoctor cleanse engine configuration that have certifications, such as SERP and AMAS, and copy them to the AddressDoctor installation directory.

For example,

• Windows: C:\AddressDoctor\5

• UNIX: /u1/addressDoctor/5

Editing the Properties FileVerify that the information in your cmxcleanse.properties files is still valid.

Related Topics:• “Editing the Properties File” on page 24

Replacing AddressDoctor 4 Library with AddressDoctor 5 LibraryYou must replace AddressDoctor 4 library with the AddressDoctor 5 library.

1. Copy the AddressDoctor 5 library from the following location:

Windows: <infamdm_install_directory>\hub\cleanse\lib\upgrade\AddressDoctorUNIX: <infamdm_install_directory>/hub/cleanse/lib/upgrade/AddressDoctor

2. Replace JADE.dll (or equivalent AddressDoctor 4 library) with the AddressDoctor 5 library at the following location:

26 Chapter 3: AddressDoctor Cleanse Engine

Page 27: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

• Windows: <infamdm_install_directory>\hub\cleanse\lib• UNIX: <infamdm_install_directory>/hub/cleanse/libFor more information, refer to the libupdate_readme.txt document available at the following location:Windows: <infamdm_install_directory>\hub\cleanse\lib\upgradeUNIX: <infamdm_install_directory>/hub/cleanse/lib/upgrade

3. Restart the application server:

a. Ensure that you are logged in with the same user name that is currently running the application server.

b. Restart the application server.

c. Check that no exceptions occur while starting the application server.

Steps to Upgrade to AddressDoctor 5 IntegrationThis section describes the upgrade process required for the MDM Hub implementation to use AddressDoctor 5.

Note: This section is applicable to users with a license for using AddressDoctor.

You must perform the following steps to upgrade to AddressDoctor 5 integration:

1. Open the cmxcleanse.properties file.This file is located at:

Windows: <infamdm_install_directory>\hub\cleanse\resourcesUNIX: <infamdm_install_directory>/hub/cleanse/resources

2. Ensure that the following AddressDoctor 5 properties are set in the cmxcleanse.properties files:

Windows:cleanse.library.addressDoctor.property.SetConfigFile=C:\infamdm\hub\cleanse\resources\AddressDoctor\5\SetConfig.xmlcleanse.library.addressDoctor.property.ParametersFile=C:\infamdm\hub\cleanse\resources\AddressDoctor\5\Parameters.xmlcleanse.library.addressDoctor.property.DefaultCorrectionType=PARAMETERS_DEFAULT

UNIX:cleanse.library.addressDoctor.property.SetConfigFile=/u1/infamdm/hub/cleanse/resources/AddressDoctor/5/SetConfig.xmlcleanse.library.addressDoctor.property.ParametersFile=/u1/infamdm/hub/cleanse/resources/AddressDoctor/5/Parameters.xmlcleanse.library.addressDoctor.property.DefaultCorrectionType=PARAMETERS_DEFAULT

3. Save and close the properties file.

4. Copy SetConfig.xml and Parameters.xml to the location specified in the cmxcleanse.properties file.

The following is a sample SetConfig.xml file:

<!DOCTYPE SetConfig SYSTEM 'SetConfig.dtd'><SetConfig> <General WriteXMLEncoding="UTF-16" WriteXMLBOM="NEVER" MaxMemoryUsageMB="600" MaxAddressObjectCount="10" MaxThreadCount="10" /> <UnlockCode>79FYL9UAXAVSR0KLV1TDC6PAQVVC3KM14FZC</UnlockCode>

Steps to Upgrade to AddressDoctor 5 Integration 27

Page 28: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

<DataBase CountryISO3="ALL" Type="BATCH_INTERACTIVE" Path="c:\addressdoctor\5" PreloadingType="NONE" />

<DataBase CountryISO3="ALL" Type="FASTCOMPLETION" Path="c:\addressdoctor\5" PreloadingType="NONE" />

<DataBase CountryISO3="ALL" Type="CERTIFIED" Path="c:\addressdoctor\5" PreloadingType="NONE" /> <DataBase CountryISO3="ALL" Type="GEOCODING" Path="c:\addressdoctor\5" PreloadingType="NONE" /> <DataBase CountryISO3="ALL" Type="SUPPLEMENTARY" Path="c:\addressdoctor\5" PreloadingType="NONE" /></SetConfig>

The following is a sample Parameters.xml file:

<?xml version="1.0" encoding="iso-8859-1"?><!DOCTYPE Parameters SYSTEM 'Parameters.dtd'><Parameters WriteXMLEncoding="UTF-16" WriteXMLBOM="NEVER"> <Process Mode="BATCH" EnrichmentGeoCoding="ON" EnrichmentCASS="ON" EnrichmentSERP="ON" EnrichmentSNA="ON" EnrichmentSupplementaryGB="ON" EnrichmentSupplementaryUS="ON" /> <Input Encoding="UTF-16" FormatType="ALL" FormatWithCountry="ON" FormatDelimiter="PIPE" /> <Result AddressElements="STANDARD" Encoding="UTF-16" CountryType="NAME_EN" FormatDelimiter="PIPE" /></Parameters>

5. Specify the AddressDoctor 5 unlock code in the configuration file, SetConfig.xml.

For more information about the SetConfig.xml file and Parameters.xml file, refer to your AddressDoctor 5 documentation.

6. Copy the AddressDoctor 5 library from the following location:

Windows: <infamdm_install_directory>\hub\cleanse\lib\upgrade\AddressDoctorUNIX: <infamdm_install_directory>/hub/cleanse/lib/upgrade/AddressDoctor

7. Replace JADE.dll (or equivalent AddressDoctor 4 library) with the AddressDoctor 5 library at the following location:

Windows: <infamdm_install_directory>\hub\cleanse\libUNIX: <infamdm_install_directory>/hub/cleanse/libFor more information, refer to the libupdate_readme.txt document available at:

Windows: <infamdm_install_directory>\hub\cleanse\lib\upgradeUNIX: <infamdm_install_directory>/hub/cleanse/lib/upgrade

8. Restart the application server.

Ensure that you are logged in with the same user name that is currently running the application server and that no exceptions occur while starting the application server.

28 Chapter 3: AddressDoctor Cleanse Engine

Page 29: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

9. Restart the Process Server.

During the Process Server initialization, you should see a message similar to the following in the terminal console:

[INFO ] com.siperian.mrm.cleanse.addressDoctor.Library: Initializing AddressDoctor510. Start the Cleanse Functions tool.

11. Obtain a write lock (Write Lock > Acquire Lock).

12. Select the AddressDoctor cleanse function.

13. Click the Refresh button.

The AddressDoctor5 cleanse function is added to the AddressDoctor cleanse functions node.

AddressDoctor 5 Fields and Process Status ValuesThe AddressDoctor 5 input fields, output fields, and process status values are significantly different from those used in AddressDoctor 4. If you have upgraded from AddressDoctor 4 to AddressDoctor 5, use the new AddressDoctor 5 field names and process status values. If you use the AddressDoctor 4 field names and process values, the MDM Hub automatically converts them to the AddressDoctor 5 names and values, but you will not benefit from the new features in AddressDoctor 5. This article provides examples and descriptions of the AddressDoctor 5 input and output fields and a comparison of the process status values to assist in the transition from AddressDoctor 4 to AddressDoctor 5.

AddressDoctor 5 Input FieldsThe input field names in AddressDoctor 5 are different than those in AddressDoctor 4. See the descriptions and examples in the following table to understand what data is handled by a particular field.

The following table lists all the AddressDoctor 5 input fields. The 'Field #' column indicates the range of fields available. For example, the available field names for AEBuilding#COMPLETE are: AEBuilding1COMPLETE, AEBuilding2COMPLETE, AEBuilding3COMPLETE, AEBuilding4COMPLETE, AEBuilding5COMPLETE, and AEBuilding6COMPLETE.

AddressDoctor 5 Input Field Name Field # Description

ACComplete - Contains the formatted address, with each formatted address line separated by a vertical bar character ('|').

AEBuilding#COMPLETE 1-6 See example 3.

AEBuilding#COMPLETE_WITH_SUBBUILDING 1-6 Contains the complete building and subbuilding information.

AEBuilding#DESCRIPTOR 1-6 See example 3.

AEBuilding#NAME 1-6 See example 3.

AEBuilding#NUMBER 1-6 Contains the building number.

AEContact#COMPLETE 1-3 See examples 1, 2, 3, and 4.

AddressDoctor 5 Fields and Process Status Values 29

Page 30: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Input Field Name Field # Description

AEContact#FIRST_NAME 1-3 See examples 1, 2, 3, and 4.

AEContact#FUNCTION 1-3 Contains the job function of the contact.

AEContact#GENDER 1-3 Contains the gender of the contact.

AEContact#LAST_NAME 1-3 See examples 1, 2, 3, and 4.

AEContact#MIDDLE_NAME 1-3 See examples 1, 2, and 3.

AEContact#NAME 1-3 See examples 1, 2, 3, and 4.

AEContact#SALUTATION 1-3 See examples 1, 2, 3, and 4.

AEContact#TITLE 1-3 Contains the job title of the contact.

AECountry#ABBREVIATION 1-3 Contains the abbreviated country name.

AECountry#ISO_NUMBER 1-3 See example 4.

AECountry#ISO2 1-3 See examples 1, 2, 3, and 4.

AECountry#ISO3 1-3 See examples 1, 2, 3, and 4.

AECountry#NAME 1-3 See examples 1, 2, 3, and 4.

AEDeliveryService#ADD_INFO 1-3 Contains additional delivery service information that passes through validation unchanged.

AEDeliveryService#COMPLETE 1-3 See example 1.

AEDeliveryService#DESCRIPTOR 1-3 See example 1.

AEDeliveryService#NUMBER 1-3 See example 1.

AEKey#RECORD_ID 1-3 -

AEKey#TRANSACTION_KEY 1-3 -

AELocality#ADD_INFO 1-6 Contains locality information that passes through validation unchanged.

AELocality#COMPLETE 1-6 See examples 1, 2, 3, and 4.

AELocality#NAME 1-6 See examples 1, 2, 3, and 4.

AELocality#SORTING_CODE 1-6 Contains a sorting code that defines where mail is sorted for localities with more than one sorting location.

AENumber#ADD_INFO 1-6 Contains additional street number information that passes through validation unchanged.

AENumber#COMPLETE 1-6 See example 3.

30 Chapter 3: AddressDoctor Cleanse Engine

Page 31: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Input Field Name Field # Description

AENumber#DESCRIPTOR 1-6 See example 3.

AENumber#NUMBER 1-6 See examples 1, 2, 3, and 4.

AEOrganization#COMPLETE 1-3 See example 4.

AEOrganization#DEPARTMENT 1-3 See example 4.

AEOrganization#DESCRIPTOR 1-3 See example 4.

AEOrganization#NAME 1-3 See example 4.

AEPostalCode#FORMATTED 1-3 See example 2.

AEPostalCode#UNFORMATTED 1-3 See example 2.

AEPostalCode#BASE 1-3 See example 2.

AEPostalCode#ADD_ON 1-3 See example 2.

AEProvince#ABBREVIATION 1-6 Contains the abbreviated province name.

AEProvince#COUNTRY_STANDARD 1-6 Contains the standard province name.

AEProvince#EXTENDED 1-6 See example 2.

AEResidue#NECESSARY 1-6 In parsing mode: contains address data that is redundant.

AEResidue#SUPERFLUOUS 1-6 In batch, certified, fast completion, or interactive mode: contains address data that is redundant.

AEStreet#ADD_INFO 1-6 Contains additional street information that passes through validation unchanged.

AEStreet#COMPLETE 1-6 See examples 1, 2, 3, and 4.

AEStreet#COMPLETE_WITH_NUMBER 1-6 See examples 1, 2, 3, and 4.

AEStreet#NAME 1-6 See examples 1, 2, 3, and 4.

AEStreet#POST_DESCRIPTOR 1-6 See examples 1, 2, 3, and 4.

AEStreet#POST_DIRECTIONAL 1-6 See example 4.

AEStreet#PRE_DESCRIPTOR 1-6 Contains street descriptors that appear before the street name.

AEStreet#PRE_DIRECTIONAL 1-6 See example 2.

AESubBuilding#COMPLETE 1-6 See example 3.

AESubBuilding#DESCRIPTOR 1-6 See example 3.

AddressDoctor 5 Fields and Process Status Values 31

Page 32: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Input Field Name Field # Description

AESubBuilding#NAME 1-6 See example 3.

AESubBuilding#NUMBER 1-6 See example 3.

ALCountrySpecificLocalityLine# 1-6 See example 5.

ALDeliveryAddressLine# 1-6 See example 5.

ALFormattedAddressLine# 1-19 See example 5.

ALRecipientLine# 1-6 See example 5.

MDMCorrectionType - This input has these possible values:- CORRECT_ONLY. In this mode, AddressDoctor

performs address standardization and validates the address components against address reference tables. This is the equivalent to Batch mode.

- CERTIFY_ONLY. In this mode, AddressDoctor validates an address according to the certification rules defined by the local postal authority. In the US, Certify_Only mode uses CASS. This is equivalent to Certified mode.

- CORRECT_THEN_CERTIFY. In this mode, AddressDoctor process the address in Batch mode and then in Certified mode.

- TRY_CERTIFY_THEN_CORRECT. In this mode, AddressDoctor processes the address in Certified mode. If issues are encountered, the address is processed again in Batch mode.

- PARAMETERS_DEFAULT. In this mode, the correction type is derived from the Parameters.xml file.

- PARSE_ONLY. In this mode, AddressDoctor parses the address, but does not correct or certify the address.

AddressDoctor 5 Output FieldsThe output field names in AddressDoctor 5 are different than those in AddressDoctor 4. See the descriptions and examples in the following table to understand what data is handled by a particular field.

The following table lists all the AddressDoctor 5 output fields. The 'Field #' column indicates the range of fields available. For example, the available field names for AEBuilding#COMPLETE are: AEBuilding1COMPLETE, AEBuilding2COMPLETE, AEBuilding3COMPLETE, AEBuilding4COMPLETE, AEBuilding5COMPLETE, and AEBuilding6COMPLETE.

AddressDoctor 5 Output Field Name Field # Description

ACComplete - Contains the formatted address, with each formatted address line separated by a vertical bar character. For example, addressline1|addressline2|addressline3.

AEBuilding#COMPLETE 1-6 See example 3.

32 Chapter 3: AddressDoctor Cleanse Engine

Page 33: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Output Field Name Field # Description

AEBuilding#COMPLETE_WITH_SUBBUILDING 1-6 Contains the complete building and subbuilding information.

AEBuilding#DESCRIPTOR 1-6 See example 3.

AEBuilding#NAME 1-6 See example 3.

AEBuilding#NUMBER 1-6 Contains the building number.

AEContact#COMPLETE 1-3 See examples 1, 2, 3, and 4.

AEContact#FIRST_NAME 1-3 See examples 1, 2, 3, and 4.

AEContact#FUNCTION 1-3 Contains the job function of the contact.

AEContact#GENDER 1-3 Contains the gender of the contact.

AEContact#LAST_NAME 1-3 See examples 1, 2, 3, and 4.

AEContact#MIDDLE_NAME 1-3 See examples 1, 2, and 3.

AEContact#NAME 1-3 See examples 1, 2, 3, and 4.

AEContact#SALUTATION 1-3 See examples 1, 2, 3, and 4.

AEContact#TITLE 1-3 Contains the title of the contact.

AECountry#ABBREVIATION 1-3 Contains the abbreviated country name.

AECountry#ISO_NUMBER 1-3 See example 4.

AECountry#ISO2 1-3 See examples 1, 2, 3, and 4.

AECountry#ISO3 1-3 See examples 1, 2, 3, and 4.

AECountry#NAME_CN 1-3 Contains the name of the country in Chinese.

AECountry#NAME_DA 1-3 Contains the name of the country in Danish.

AECountry#NAME_DE 1-3 Contains the name of the country in German.

AECountry#NAME_EN 1-3 Contains the name of the country in English. See examples 1, 2, 3, and 4.

AECountry#NAME_ES 1-3 Contains the name of the country in Spanish.

AECountry#NAME_FI 1-3 Contains the name of the country in Finnish.

AECountry#NAME_FR 1-3 Contains the name of the country in French.

AECountry#NAME_GR 1-3 Contains the name of the country in Greek.

AECountry#NAME_HU 1-3 Contains the name of the country in Hungarian.

AddressDoctor 5 Fields and Process Status Values 33

Page 34: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Output Field Name Field # Description

AECountry#NAME_IT 1-3 Contains the name of the country in Italian.

AECountry#NAME_JP 1-3 Contains the name of the country in Japanese.

AECountry#NAME_KR 1-3 Contains the name of the country in Korean.

AECountry#NAME_NL 1-3 Contains the name of the country in Dutch.

AECountry#NAME_PL 1-3 Contains the name of the country in Polish.

AECountry#NAME_PT 1-3 Contains the name of the country in Portuguese.

AECountry#NAME_RU 1-3 Contains the name of the country in Russian.

AECountry#NAME_SA 1-3 Contains the name of the country in Arabic.

AECountry#NAME_SE 1-3 Contains the name of the country in Swedish.

AEDeliveryService#ADD_INFO 1-3 Contains additional delivery service information.

AEDeliveryService#COMPLETE 1-3 See example 1.

AEDeliveryService#DESCRIPTOR 1-3 See example 1.

AEDeliveryService#NUMBER 1-3 See example 1.

AEKey#RECORD_ID 1-3 Contains the record ID.

AEKey#TRANSACTION_KEY 1-3 Contains the transaction key.

AELocality#ADD_INFO 1-6 Contains portions of the locality that could not be validated against reference data.

AELocality#COMPLETE 1-6 See examples 1, 2, 3, and 4.

AELocality#NAME 1-6 See examples 1, 2, 3, and 4.

AELocality#SORTING_CODE 1-6 Contains a sorting code defining where mail is sorted for localities with more than one sorting location.

AENumber#ADD_INFO 1-6 Contains portions of the street number that could not be validated against reference data.

AENumber#COMPLETE 1-6 See example 3.

AENumber#DESCRIPTOR 1-6 See example 3.

AENumber#NUMBER 1-6 See examples 1, 2, 3, and 4.

AEOrganization#COMPLETE 1-3 See example 4.

AEOrganization#DEPARTMENT 1-3 See example 4.

34 Chapter 3: AddressDoctor Cleanse Engine

Page 35: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Output Field Name Field # Description

AEOrganization#DESCRIPTOR 1-3 See example 4.

AEOrganization#NAME 1-3 See example 4.

AEPostalCode#FORMATTED 1-3 See example 2.

AEPostalCode#UNFORMATTED 1-3 See example 2.

AEPostalCode#Base 1-3 See example 2.

AEPostalCode#ADD_ON 1-3 See example 2.

AEProvince#ABBREVIATION 1-6 Contains the abbreviated province name.

AEProvince#COUNTRY_STANDARD 1-6 Contains the standard province name.

AEProvince#EXTENDED 1-6 See example 2.

AEResidue#NECESSARY 1-6 In parsing mode: contains address data that is redundant.

AEResidue#SUPERFLUOUS 1-6 In batch, certified, fast completion, or interactive mode: contains address data that is redundant.

AEResidue#UNRECOGNIZED 1-6 Contains data that cannot be parsed to an address port.

AEStreet#ADD_INFO 1-6 Contains portions of the street information that could not be validated against reference data.

AEStreet#COMPLETE 1-6 See examples 1, 2, 3, and 4.

AEStreet#COMPLETE_WITH_NUMBER 1-6 See examples 1, 2, 3, and 4.

AEStreet#NAME 1-6 See examples 1, 2, 3, and 4.

AEStreet#POST_DESCRIPTOR 1-6 See examples 1, 2, 3, and 4.

AEStreet#POST_DIRECTIONAL 1-6 See example 4.

AEStreet#PRE_DESCRIPTOR 1-6 Contains street descriptors that appear before the street name.

AEStreet#PRE_DIRECTIONAL 1-6 See example 2.

AESubBuilding#COMPLETE 1-6 See example 3.

AESubBuilding#DESCRIPTOR 1-6 See example 3.

AESubBuilding#NAME 1-6 See example 3.

AESubBuilding#NUMBER 1-6 See example 3.

ALCountrySpecificLocalityLine# 1-6 See example 5.

AddressDoctor 5 Fields and Process Status Values 35

Page 36: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Output Field Name Field # Description

ALDeliveryAddressLine# 1-6 See example 5.

ALFormattedAddressLine# 1-19 See example 5.

ALRecipientLine# 1-6 See example 5.

EDCASSStatus - Indicates if the address contains enough data for USPS Coding Accuracy Support System (CASS) certification.

EDGeoCodingStatus - Indicates the level of accuracy of the Geocode value.

EDSERPStatus - Indicates if the address contains enough data for Canada Post Software Evaluation and Recognition Program (SERP) certification.

EDSNAStatus - Indicates if the address contains enough data for France's La Post National Address Management Service (SNA) certification.

EDSupplementaryGBStatus - Indicates whether GB country-specific output is available.

EDSupplementaryUSStatus - Indicates whether US country-specific output is available.

EECASSBARCODE - Contains the USPS barcode number for the address.

EECASSCARRIER_ROUTE - Contains the USPS carrier route for a US address.

EECASSCONGRESSIONAL_DISTRICT - Contains the congressional district based on the United States Postal Service (USPS) Zip+4 code.

EECASSDEFAULT_FLAG - Indicates if a US address matches a high-rise exact, high-rise default, or rural route default address in the reference data.

EECASSDELIVERY_POINT - Contains a unique Delivery Point Code (DPC) assigned to each USPS address in a Zip+4 code area.

EECASSDELIVERY_POINT_CHECK_DIGIT - Contains a number between 0 and 9 that, when added with the digits in the USPS Zip+4 and DPC, create a sum divisible by 10.

EECASSDPV_CMRA - Indicates if the address is for a USPS Commercial Mail Receiving Agent (CMRA), such as a post office box, and not the physical location of a business or residence.

EECASSDPV_CONFIRMATION - Contains a code indicating to what degree the USPS DPC value is valid.

36 Chapter 3: AddressDoctor Cleanse Engine

Page 37: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Output Field Name Field # Description

EECASSDPV_FALSE_POSITIVE - Indicates whether an address was generated from encrypted USPS reference data.

EECASSDPV_FOOTNOTE_1 - Indicates whether the address matches the USPS Zip+4 code data.

EECASSDPV_FOOTNOTE_2 - Contains a DPV status code.

EECASSDPV_FOOTNOTE_3 - Indicates whether an address is a USPS CMRA.

EECASSDPV_FOOTNOTE_COMPLETE - Contains the combined data for the DPV footnotes 1, 2, and 3.

EECASSDSF2_NOSTATS_INDICATOR - Indicates whether a USPS address is valid but undeliverable.

EECASSDSF2_VACANT_INDICATOR - Indicates whether a USPS address is inactive.

EECASSERRORCODE - Contains the CASS error code.

EECASSEWS_RETURNCODE - Indicates whether an address is in the USPS Early Warning System (EWS) list of new addresses that are not yet referenced to a ZIP+4 level.

EECASSHIGHRISE_DEFAULT - Indicates whether a USPS address matches a high-rise record in the address reference data and does not contain a unit identifier.

EECASSHIGHRISE_EXACT - Indicates whether a USPS address matches a high-rise record in the address reference data and also contains a unit identifier.

EECASSLACS - Indicates whether a United States address is present in the USPS Locatable Address Conversion Service (LACS) table.

EECASSLACSLINK_INDICATOR - Indicates if the address validator checks the address against the USPS LACS reference database.

EECASSLACSLINK_RETURNCODE - Indicates the degree to which the input address matches USPS LACS data and whether the address validation process updated the address.

EECASSRECORDTYPE - Contains additional information about the deliverable status of non-DPC USPS addresses.

EECASSRURALROUTE_DEFAULT - Indicates if the address is a valid rural route but exact data is unavailable.

EECASSRURALROUTE_EXACT - Indicates if the address matches a rural route address in the USPS address reference data set.

EECASSSUITELINK_RETURNCODE - Identifies high-rise business addresses in the United States that lack suite identification information.

AddressDoctor 5 Fields and Process Status Values 37

Page 38: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Output Field Name Field # Description

EECASSZIPMOVE_RETURNCODE - Indicates if the USPS has recently changed the ZIP+4 Code assigned to the address.

EEGeoCodingCOMPLETE - Contains the complete geocode coordinates for the output address.

EEGeoCodingLATITUDE - Contains the latitude coordinate of the address.

EEGeoCodingLONGITUDE - Contains the longitude coordinate of the address.

EEGeoCodingLAT_LONG_UNIT - Contains the latitude and longitude unit of measure.

EESERPCATEGORY - Contains the SERP Certification status code.

EESNACATEGORY - Contains the SNA Certification status code.

EESupplementaryGBDELIVERY_POINT_ SUFFIXES

- Contains the two-character suffix assigned to a mailbox in a UK post code area (Royal Mail).

EESupplementaryUSCBSA_ID - Contains a USPS Core-Based Statistical Area (CBSA) identification number. A CBSA identifies an urban area with a population greater than 10,000.

EESupplementaryUSCOUNTY_FIPS_CODE - Contains the US Federal Information Processing Standard (FIPS) Code that identifies a county or county equivalent in the United States and United States possessions.

EESupplementaryUSFINANCE_NUMBER - Contains the code assigned to United States post offices and other postal facilities to enable collection of cost and statistical data.

EESupplementaryUSMSA_ID - Contains the USPS Metropolitan Statistical Area identification number. This number identifies an urban area with a population greater than 50,000.

EESupplementaryUSRECORD_TYPE - Contains a single-character code that describes the type of mailbox.

EESupplementaryUSSTATE_FIPS_CODE - Contains the FIPS Code that identifies a state or state equivalent in the United States and United States possessions.

MDMLastError - The MDM Last Error output. Not associated with the AddressDoctor result.

RDElementInputStatus - Indicates the similarities between input address data and the address reference data.

RDElementRelevance - Indicates if an address element is required for postal delivery.

RDElementResultStatus - Contains a description of changes made to input address data elements during address validation.

38 Chapter 3: AddressDoctor Cleanse Engine

Page 39: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 5 Output Field Name Field # Description

RDMailabilityScore - Contains a single digit that represents the likelihood of successful delivery to the validated address, based on overall validation results.

RDResultPercentage - Indicates the percentage likelihood of successful delivery to the validated address.

RPCount - Contains the number number of results available.

RPCountOverflow - Indicates whether there are more results available than the 20 results that were returned.

RPCountryISO3 - Contains the ISO3 country code.

RPModeUsed - Indicates which mode was used to process the address data.

RPPreferredLanguage - Identifies which language should be returned.

RPPreferredScript - Identifies which alphabet should be returned.

RPProcessStatus - Contains the process status value.

Address ExamplesThe following examples show how an address is broken down into specific address elements in the AddressDoctor 5 address element fields.

Example 1: Canadian AddressThe sample Canadian address used in the example can be broken down into address element fields, as shown in the following illustration:

Mr. John Edward Doe1425 10th AvenuePO Box 4001Victoria BC V8X 3X4Canada

AddressDoctor 5 Fields and Process Status Values 39

Page 40: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Example 2: US AddressThe sample US address used in the example can be broken down into address element fields, as shown in the following illustration:

Mr. John Edward Doe118 West Aaron SquareAaronsburg PA 16820-9407United States

40 Chapter 3: AddressDoctor Cleanse Engine

Page 41: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Example 3: Hong Kong Address with Buildings and SubbuildingsThe sample Hong Kong address used in the example can be broken down into address element fields, as shown in the following illustration:

Mr. John Edward TsangFloor 12, Apt AFung Wah Estate, Hiu Fung TowerNo. 11 North StreetKennedy TownHong Kong IslandHKG

Example 4: UK Business AddressThe sample UK business address used in the example can be broken down into address element fields, as shown in the following illustration:

DEF Inc.RDN GroupMr. John Doe38A St. James's Street EastLondonE17 7PEUnited Kingdom

AddressDoctor 5 Fields and Process Status Values 41

Page 42: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Example 5: Using Address Line Fields to Simplify Address EntryInputting address data into address line fields is much easier than inputting the address into individual address elements. The following example shows how an address is broken down into address lines in the AddressDoctor 5 address line fields. The recipient, delivery address, and country-specific locality field names indicate the type of address data contained in the field. Use the formatted address line fields when the address data entered into a line field is not constant, such as when there are a combination of residential and business addresses. The type of address data in a given address line varies by country. Ensure that the address data is entered into the appropriate address line.

John Doe123 Main St NW STE 12Anytown NY 12345

42 Chapter 3: AddressDoctor Cleanse Engine

Page 43: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Differences Between AddressDoctor 4 and AddressDoctor 5 Process Status Values

The process status values indicate the reliability of the output address, the address verification result, and if the input address was corrected. Some process status values have changed in AddressDoctor 5 and some process status values were introduced to provide more detailed information.

The following table compares the AddressDoctor 4 and AddressDoctor 5 process status values.

AddressDoctor 4 Process Status Value

AddressDoctor 5 Process Status Value

Description

V - Verified. The input data is correct.

- V4 Verified. The input data is correct. All relevant elements were checked and the input data matched the reference data.

- V3 Verified. The input data is correct but some or all elements were standardized or the input contains outdated names or exonyms.

- V2 Verified. The input data is correct but some elements could not be verified because of incomplete reference data.

- V1 Verified. The input data is correct but incorrect element user standardization has deteriorated deliverability. For example, the postcode length chosen is too short. Not set by validation.

C - Corrected.

- C4 Corrected. All relevant elements have been checked.

- C3 Corrected. However, some elements could not be checked.

- C2 Corrected. However, the delivery status is unclear due to a lack of reference data.

- C1 Corrected. However, the delivery status is unclear because user standardization is incorrect. Not set by validation.

P3 - Data cannot be corrected, but likely to be deliverable.

- I4 Data could not be corrected, but is likely to be deliverable. An address element does not match the single reference value available. For example, the house number is incorrect and there is one house number in the reference data.

- I3 Data could not be corrected, but is likely to be deliverable. An address element does not match the multiple reference values available. For example, the house number is incorrect and there is more than one house number in the reference data.

P2 I2 Data could not be corrected, but there is a slim chance that the address is deliverable.

P1 I1 Data cannot be corrected and unlikely to be deliverable.

AddressDoctor 5 Fields and Process Status Values 43

Page 44: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 4 Process Status Value

AddressDoctor 5 Process Status Value

Description

- N5 Validation Error. Validation was not performed because the reference database is too old. Contact AddressDoctor to get updated reference data.

N4 - Validation method not yet called after parsing operation.

- N4 Validation Error. Validation was not performed because the reference database is corrupted or in the wrong format.

N3 N3 Validation Error. Validation was not performed because the country could not be unlocked.

N2 N2 Validation Error. Validation was not performed because the required reference database is not available.

N1 N1 Validation Error. Validation was not performed because the country was not recognized.

Q3 Q3 FastCompletion Status. The suggested address is complete.

Q2 Q2 FastCompletion Status. The suggested address is complete. However, elements were deleted or added from the input.

Q1 Q1 FastCompletion Status. The suggested address is not complete. Enter more information.

Q0 Q0 FastCompletion Status. Suggestions cannot be generated. The input data is insufficient.

- S4 Parsed perfectly.

- S3 Parsed with multiple results.

- S2 Parsed with errors. Elements changed position.

- S1 Parse error. Input format mismatch.

- RB Country recognized from abbreviation.

- RA Country recognized from ForceCountryISO3 setting.

- R9 Country recognized from DefaultCountryISO3 setting.

- R8 Country recognized from name without errors.

- R7 Country recognized from name with errors.

- R6 Country recognized from territory.

- R5 Country recognized from province.

- R4 Country recognized from major town.

44 Chapter 3: AddressDoctor Cleanse Engine

Page 45: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

AddressDoctor 4 Process Status Value

AddressDoctor 5 Process Status Value

Description

- R3 Country recognized from format.

- R2 Country recognized from script.

- R1 Country not recognized. Multiple matches.

- R0 Country not recognized.

Configuring the JVM SettingsIf you use AddressDoctor as your cleanse engine to run any batch process, you must ensure that the stack size for the JVM is sufficient.

If an incorrect stack size is set for your application server, AddressDoctor might cause an initialization failure when you start the Process Server. Also, it might throw an exception during a cleanse operation when the Process Server is up and running.

Setting the JVM Size for WebSphere on Windows/UNIXTo set the JVM size for WebSphere:

1. Open the WebSphere Console.

2. Navigate to Servers> Application Server> <Your Server>>Process Definition> Java Virtual Machine.

3. Add the following parameters to the Generic JVM Arguments: -Xss1024k - Parameter to set JVM stack size of 1 MB-Xms512m - Parameter to set JVM heap with a start (minimum) size of 512 MB-Xmx1024m - Parameter to set JVM heap to allow growth to maximum of 1 GB

4. If you use AddressDoctor 5.4 or higher, add the following parameter to the Generic JVM Arguments: -Xmso2048k - Parameter to increase the operating system thread stack size

5. Save the configuration.

6. Restart the WebSphere application server profile that runs the Process Server.

Setting the JVM Size for WebLogic Server on WindowsTo set the JVM size for the WebLogic server on Windows:

1. Go to your WebLogic home directory.

2. Open setDomainEnv.cmd in a text editor.

3. Set the MEM_ARGS variable as follows:

set MEM_ARGS=%MEM_ARGS% -Xmx1024m4. Save and close the setDomainEnv.cmd file.

Configuring the JVM Settings 45

Page 46: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Setting the JVM Size for WebLogic Server on UNIXTo set the JVM size for WebLogic server on HP-UX and Solaris:

1. Go to your WebLogic home directory.

2. Open setDomainEnv.sh in a text editor.

3. Set the MEM_ARGS variable as follows:

set MEM_ARGS=$MEM_ARGS -Xmx1024m4. Save and close the setDomainEnv.sh file.

Setting the JVM Size for JBoss on WindowsTo set the JVM size for the JBoss server on Windows:

1. Go to your JBoss home directory.

2. Navigate to the bin directory.

3. Open run.bat in a text editor.

4. Set the JAVA_OPTS variable as follows: JAVA_OPTS=%JAVA_OPTS% -Xmx1024m

5. Save and close the run.bat file.

Setting the JVM Size for JBoss on UNIXTo set the JVM size for JBoss on UNIX:

1. Go to your JBoss installation directory.

2. Navigate to the bin directory.

3. Open run.sh in a text editor.

4. Set the JAVA_OPTS variable as follows: JAVA_OPTS="$JAVA_OPTS -Xmx1024m

5. Save and close the run.sh file.

After you make configuration changes, restart your Informatica MDM Hub and Process Servers.

46 Chapter 3: AddressDoctor Cleanse Engine

Page 47: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

C H A P T E R 4

FirstLogic Direct Cleanse EngineThis chapter includes the following topics:

• FirstLogic Direct Cleanse Engine Overview, 47

• About FirstLogic Direct Integration, 47

• Installing the Components, 48

• Configuring FirstLogic Direct, 48

• Configuring Informatica MDM Hub to Use the Adapter, 48

• Using Your FirstLogic Direct Library, 49

FirstLogic Direct Cleanse Engine OverviewIntegration between Informatica MDM Hub and the FirstLogic Direct cleanse engine takes place through the Informatica MDM Hub FirstLogicDirect Adapter. This adapter is an optional component.

This chapter explains how to configure your Informatica MDM Hub system to use the FirstLogicDirect Adapter and the FirstLogic Direct cleanse engine on both the Windows and UNIX platforms. The information in this chapter pertains to both platforms unless otherwise specified.

This chapter assumes that you are knowledgeable about configuration and using the FirstLogic Direct software. To learn more about FirstLogic Direct, see your FirstLogic Direct documentation.

About FirstLogic Direct IntegrationInformatica MDM Hub integrates with FirstLogic Direct by treating FirstLogic Direct transaction-type projects as cleanse functions.

The process of integration with FirstLogic Direct involves both Informatica MDM Hub and FirstLogic Direct.

To integrate the FirstLogic Direct cleanse engine with Process Server:

1. Install the components of the Process Server and the cleanse engine.

2. Configure FirstLogic Direct.

3. Configure Informatica MDM Hub to use the adapter.

Note: You must obtain a license (with cleanse adapter and Business Objects DQ XI enabled) from Informatica.

47

Page 48: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Once these steps have been successfully completed, you are ready to use the FirstLogic Direct functions in the Process Server.

Installing the ComponentsBefore you configure the adapter, ensure that you have installed the components:

1. Install the application server you intend to use for the Process Server.

2. Install FirstLogic Direct.

3. Install the Process Server as explained in the Informatica MDM Hub Installation Guide. Process Server and FirstLogic Direct can be on the same machine or different machines.

For more information about installing the components, see the Informatica MDM Hub Installation Guide.

Configuring FirstLogic DirectThis table describes the actions required to configure your Business Objects DQ XI to work with the Informatica MDM Hub.

File or Component Name

Action Notes

System environment variable JAVA_HOME

Ensure that this variable:Points to the correct (supported) Java version (see the Release Notes compatibility information).Is found on the PATH variable before other Java versions.

FirstLogic Direct must specify its own Java version.

Business Objects DQ XI server

Start the FirstLogic Direct server.Windows: Go to:Start > Control Panel > Administrative Tools > Services.UNIX:Go to <dqxi_install_dir>/bin.Type this command to start the dqxi service:execute start_dqxiserver.sh

Configuring Informatica MDM Hub to Use the Adapter

In order to integrate FirstLogic Direct with Informatica MDM Hub, you must configure Informatica MDM Hub for the adapter.

48 Chapter 4: FirstLogic Direct Cleanse Engine

Page 49: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

All parameters are set during the Process Server installation. and can be modified inside FLD_default_config.txt file. This file is located in the following directory:

• Windows:<infamdm_install_directory>\cleanse\resources\firstLogicDirect

• UNIX:<infamdm_install_directory>/cleanse/resources/firstLogicDirect

When the parameters are set, Informatica MDM Hub imports the project list and adds each transaction type as an Informatica MDM Hub cleanse function.

On every load of the library in the Hub Console or on every refresh of the FirstLogic DQXI SDK cleanse functions, the FirstLogicDirect Adapter:

• Does a lookup to the FirstLogic Direct server.

• Retrieves all transactional projects, with their input and output parameters.

• Displays those projects as FirstLogic Direct functions in the Informatica MDM Hub cleanse tool.

Using Your FirstLogic Direct LibraryYou can use your FirstLogic Direct transactional project in Informatica MDM Hub’s cleanse tool just as you would use any other cleanse function. The FirstLogicDirect Adapter lists all your transactional projects in Informatica MDM Hub’s cleanse tool. When you create new projects in FirstLogic Direct, or modify a project’s inputs or outputs, you must refresh the view in Informatica MDM Hub’s cleanse tool to see the updates.

About Transactional ModeIn order to use FirstLogic Direct in transactional mode, the project must be of type Transaction. An example of a transactional project is in the FirstLogic Direct blueprints:

By default, these files are located in the following directory:

• Windows:<dqxi_Install_Dir>\repository\configuration_rules\projects\blueprints\transactional_address_data_cleanse_us.xml

• UNIX:<dqxi_Install_Dir>/repository/configuration_rules/projects/blueprints/transactional_address_data_cleanse_us.xml

By default, this project uses inputs Address Line 1 through 4, and outputs as delivery-address, city, region, postcode, and so on. You can create additional output fields by adding fields to the SOAP Writer component in the Output Field Options group. Specific values can be derived from the Output Field Options of the USA Multi-line component.

About the Transactional Mode SampleYour distribution includes a sample FirstLogic Direct project for the Process Server. This was created by customizing the FirstLogic Direct “blueprint”:

Using Your FirstLogic Direct Library 49

Page 50: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

The blueprint is found in the following directory:

• Windows:<dqxi_Install_Dir>\repository\configuration_rules\projects\blueprints\trans_address_data_cleanse_usa.xml

• UNIX:<dqxi_Install_Dir>/repository/configuration_rules/projects/blueprints/trans_address_data_cleanse_usa.xml

The sample transactional project is configured to input up to four (4) fields of address data and up to eight (8) fields of non-address data in an XML string, cleanse the address, name, firm, SSN (U.S.only), date, phone, and email data, and output discrete cleansed fields in an XML string.

50 Chapter 4: FirstLogic Direct Cleanse Engine

Page 51: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

C H A P T E R 5

Trillium Director Cleanse EngineThis chapter includes the following topics:

• Trillium Director Cleanse Engine Overview, 51

• About Trillium Director Integration, 51

• Before You Install, 52

• Configuring Trillium Director and the Cleanse Match Server, 52

• Upgrading Trillium Director, 56

• Using Trillium on a Remote Server, 56

• Configuring Trillium Director for Multithreading, 57

• Setting the Threading Pool, 58

• Increasing the Number of Network Connection Retries, 58

Trillium Director Cleanse Engine OverviewIntegration between Informatica MDM Hub and the Trillium Director cleanse engine occurs through the Trillium Director Adapter. This adapter is an optional component.

This chapter explains how to configure your Informatica MDM Hub system to use the Trillium Director Adapter and the Trillium Director cleanse engine. This chapter assumes that you are knowledgeable about configuring and using the Trillium Director software. To learn more about Trillium Director, see your Trillium Director documentation.

Note: Ensure that you have the latest Trillium patches installed before you work with the Trillium Director Adapter. Contact Informatica Global Customer Support for more information.

About Trillium Director IntegrationInformatica MDM Hub integrates with Trillium Director by treating Trillium Director projects as cleanse functions. Informatica MDM Hub uses a Trillium Director configuration file to determine the Trillium Director

51

Page 52: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

projects that are available for integration into the Hub Console library. The default configuration files are located in the following directory:

• Windows:<infamdm_install_directory>\cleanse\resources\Trillium\samples\director\

• UNIX:<infamdm_install_directory>/cleanse/resources/Trillium/samples/director/

The process of integrating with Trillium Director involves installing and configuring Trillium Director for use with Informatica MDM Hub.

To integrate the Trillium Director cleanse engine with Process Server:

1. Ensure that the prerequisites for Trillium Director Adapter are met.

2. Install Trillium Director. To learn more, see your Trillium Director documentation.

Important: We recommend that you install the Process Server and Trillium Director on the same machine for improved performance.

3. Install your application server. To learn more, see your application server documentation.

4. Install Process Server as explained in the Informatica MDM Hub Installation Guide.

5. Configure Trillium Director to use the Informatica MDM Hub.

6. Check your installation and configuration.

Before You InstallBefore you install:

1. Ensure that your Informatica MDM Hub license has the cleanse adapter and the Trillium Director Adapter enabled. Contact Informatica Global Customer Support to obtain this license.

2. Ensure that the following Trillium Director components are installed on the same machine as the Process Server:

• Trillium Director Server

• Trillium Director data files for the countries with which you are working. Contact Trillium for these files.

Configuring Trillium Director and the Cleanse Match Server

In order to successfully integrate Trillium Director 11 or later with the Informatica MDM Hub, you must configure both Trillium Director 11 or later and the Process Server to work together.

1. Set the Trillium Director port, so that it does not conflict with any other port that is used on the machine. To learn more, see your Trillium Director documentation.

2. Go to the following directory:

52 Chapter 5: Trillium Director Cleanse Engine

Page 53: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

On Windows:<Cleanse_Install_Directory>\resources\TrilliumDirector11\samples\director\

On UNIX:<Cleanse_Install_Directory>/resources/TrilliumDirector11/samples/director/

Check to ensure you have the following files:

• td11_default_config_Global.txt

• td11_default_config_US_detail.txt

• td11_default_config_US_summary.txt

These files define the Trillium Director projects you use with the Process Server.

If you do not have these files, contact Informatica Global Customer Support.

Note: You can choose to create your own configuration files.

3. Open the cmxcleanse.properties file for editing:

On Windows:<Cleanse_Match_Server_Install_Dir>\resources\cmxcleanse.properties

On UNIX:<Cleanse_Match_Server_Install_Directory>/resources/cmxcleanse.properties

4. Check the following values for your Trillium Director version.

Sample values for Trillium Director 11:cleanse.library.trilliumDir11.property.config.file.1=<infamdm_install_directory>/hub /cleanse/resources/TrilliumDirector11/samples/director/td11_default_config_Global.txtcleanse.library.trilliumDir11.property.config.file.2=<infamdm_install_directory>/hub /cleanse/resources/TrilliumDirector11/samples/director/td11_default_config_US_detail.txtcleanse.library.trilliumDir11.property.config.file.3=<infamdm_install_directory>/hub /cleanse/resources/TrilliumDirector11/samples/director/td11_default_config_US_summary.txt

5. Save and close the properties file.

6. Connect to the database as the schema owner and run the following statement to enable Trillium Director:

update C_REPOS_CL_FUNCTION_LIB set DISPLAY_IND = 1 where function_lib_name = 'Trillium Director Version 11'

7. Add the Trillium Director bin or bin64 directory path to the CLASSPATH and the PATH environment variables:

• On 32-bit Windows:<Trillium Director install directory>\TSQuality\<Trillium version>\Software\bin

• On 64-bit Windows:Trillium Director install directory\TSQuality\<Trillium version>\Software\bin64

• On UNIX:/apps/oracle/TrilliumSoftware/tsq11r5s/Software/bin

8. Also check that the TRILLDIRPORT and TRILLDIRADDR environment variables are set to an appropriate value. For example:

TRILLDIRADDR=localhostTRILLDIRPORT=14445

9. If you have not done so already, start the Trillium Director and Trillium Cleansing Server services.

10. Start the Trillium Director and Trillium Cleansing Server services.

Configuring Trillium Director and the Cleanse Match Server 53

Page 54: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Testing Your Trillium Director ConfigurationTo test your Trillium Director configuration:

1. Go to the Informatica MDM Hub console.

2. Choose the Cleansing tool from the workbench.

3. Choose cleansing functions.

4. Refresh the Trillium Director library by using the Refresh tab on the right-hand side in the Hub Console. You should see your Trillium Director functions.

If you do not see those functions, check to ensure that you have successfully completed the steps in previous configuration sections.

You can now use these functions to make a mapping and cleanse your data.

Sample Configuration FilesUse the configuration files to configure information such as the function name, server name, and the input and output parameters.

You can use the sample configuration files, or create your own configuration files. If you use a sample configuration file, update the function name, system ID, and server name for your environment. Informatica supplies the following sample configuration files that work with Trillium Director 11 and later:td11_default_config_Global.txt

Provides output fields for international addresses.

td11_default_config_US_detail.txt

Provides output fields for U.S. addresses.

td11_default_config_US_summary.txt

Provides the most common output fields for U.S. addresses.

The sample configuration files are in the following locations:

• On Windows:<infamdm installation directory>\hub\cleanse\resources\TrilliumDirector11\samples\director\

• On UNIX:<infamdm installation directory>/hub/cleanse/resources/TrilliumDirector11/samples/director/

The following information is in the sample configuration files:

• function name

• function description

• input parameters

• output parameters

• update parameters

54 Chapter 5: Trillium Director Cleanse Engine

Page 55: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

The following table describes the entries in this configuration file:

Name Value

TD_FUNCT_NAME The name of your Trillium Director function.

TD_FUNCT_DESCR A description of the function. This appears when you look at the available functions using the Cleansing Tool.

TD_SYSTEM_ID The Trillium Director system ID. This defines which Trillium Director process to use.

TD_SERVER_NAME The name of the server process within the Trillium Director system indicated in TD_SYSTEM_ID. This is the name of the Trillium Director server that is running the project you are using. This is the Cleanser instance where your project was deployed.

TD_INPUT_PARAM_* This is a group of entries, one for each input parameter for the function. Add or remove entries as necessary.

TD_OUTPUT_PARAM_* This is a group of entries, one for each output parameter for the function. Add or remove entries as necessary.

TD_UPD_PARAM_* This is a group of entries, one for each update parameter for the function. Add or remove entries as necessary.

The following example shows a sample configuration file for Trillium Director 11 and later:

TD_FUNCT_NAME = TrilliumDirectorGlobalTD_FUNCT_DESCR = Trillium Director adapter

TD_SYSTEM_ID = GTD_SERVER_NAME = Cleanser

TD_INPUT_PARAM_0 = MRMRSTD_INPUT_PARAM_1 = BUSINESSNAMETD_INPUT_PARAM_2 = StreetAddressTD_INPUT_PARAM_3 = StreetAddress2TD_INPUT_PARAM_4 = StreetAddress3TD_INPUT_PARAM_5 = StreetAddress4TD_INPUT_PARAM_6 = CityTD_INPUT_PARAM_7 = StateTD_INPUT_PARAM_8 = PostalCodeTD_INPUT_PARAM_9 = Country

-- Derived Global Output Fields (country specific output fields also available)

TD_OUTPUT_PARAM_0 = DR_HOUSE_NUMBER1TD_OUTPUT_PARAM_1 = DR_STREET_NAMETD_OUTPUT_PARAM_2 = DR_HOUSE_NUMBER2TD_OUTPUT_PARAM_3 = DR_CITY_NAMETD_OUTPUT_PARAM_4 = DR_REGION_NAMETD_OUTPUT_PARAM_5 = DR_POSTAL_CODETD_OUTPUT_PARAM_6 = DR_ADDRESSTD_OUTPUT_PARAM_7 = DR_COUNTRY_NAMETD_OUTPUT_PARAM_8 = DR_ADDR_MAIL

Configuring Trillium Director and the Cleanse Match Server 55

Page 56: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Upgrading Trillium DirectorYou can upgrade the Trillium Director cleanse engine if required.Note: When you upgrade the Informatica MDM Hub installation, it is not required to upgrade Trillium Director.

1. If upgrading to Trillium 11 or later versions, obtain an updated license file.

2. Edit the Hub Cleanse Match properties file.

Using Trillium on a Remote ServerYou must configure Trillium Director before it can run on a remote server.

1. Copy the Trillium Director files from the remote server to the MDM Hub Process Server.

• The following table shows the files and server paths for Windows 32-bit systems:

Files/Paths Description

Files TGenClient.dllxerces-c*.dll

Remote server path <Trillium Director install directory>\TSQuality\<Trillium version>\Software\bin

Process Server path <infamdm install directory>\hub\cleanse\lib

• The following table shows the files and server paths for Windows 64-bit systems:

Files/Paths Description

Files TGenClient.dllxerces-c*.dlltriTGenClientLibrary.dll

Remote server path <Trillium Director install directory>\TSQuality\<Trillium version>\Software\bin64

Process Server path <infamdm install directory>\hub\cleanse\lib

56 Chapter 5: Trillium Director Cleanse Engine

Page 57: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

• The following table shows the files and server paths for UNIX systems:

Files/Paths Description

Files libTGenClientLibrary.solibexerces-c*.so

Remote server path <Trillium Director install directory>/TSQuality/<Trillium version>/Software/bin

Process Server path <infamdm install directory>/hub/cleanse/lib

2. Ensure all the file permissions are set correctly for user and group:

• Add <infamdm_install_directory>\cleanse\lib to the LIBPATH or LD_LIBRARY_PATH environment variable.

Note: If using the WebSphere application server, do not put the LIBPATH or LD_LIBRARY_PATH environment variable into the classpath of JVM. For example, Application servers > server1 > Process Definition > Java Virtual Machine: Classpath variable.

• Ensure the following Trillium environment variables are set correctly:TRILLDIRADDR=<trillium_server_machine_address>TRILLDIRPORT=<Trillium server port>TRILLCONFIG=<Trillium Director install directory>\TSQuality\Trillium version\director_proj\settings\TrilXML.cfgTS_CONFIG: <Trillium Director install directory>\TSQuality\<Trillium version>\director_proj\settings\Config.tblTS_QUALITY=<Trillium Director install directory>\TSQuality\<Trillium version>\Software

Configuring Trillium Director for MultithreadingIf you have the Trillium Director installed on a machine with more than one processor, you can use multithreading to take advantage of your hardware.

To configure your Trillium Director for performance:

• Turn on caching for performance benefits.

• By default, Trillium Director is set to be single-threaded. Trillium Director reverts to this default if you set an invalid value for the number of threads, such as a number equal to or less than 0 or a very large number. Set the number of threads to equal the number of CPUs in the machine where Trillium Director is installed.

You can confirm the thread settings by turning debug on. This shows you threads that are spawned and other information. To learn more, see your Trillium Director documentation.

Important: Turn debug off when you are in a production environment. Debug can have a significant negative effect on performance.

Configuring Trillium Director for Multithreading 57

Page 58: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Setting the Threading PoolWhen the Trillium adapter works with large data sets, it uses thread-pool connections in the application server. These threads are shared by all the jobs that the server handles. The thread-pool level is designated in the following property: cleanse.library.trilliumDir.property.td.project.pool.MaxActive=nn where nn is the same or slightly higher than the total thread count of all jobs running in parallel.

The default setting for the MaxActive property allows as many connections as are requested. To restrict the number of connections, reset the MaxActive property value.

Increasing the Number of Network Connection Retries

When network traffic is heavy and a high number of cleanse threads are being used, the network connection for a thread between the Informatica MDM Hub and the Trillium server can time out while a record is being processed. The server automatically tries to restart the connection five times, and issues an error message after the fifth try if the connection attempt fails.

Note: The batch cleanse still runs through, with the record marked as rejected.

The number of restart tries is determined by the following property, which can be added to the cmxcleanse.properties file:

cleanse.library.trilliumDir.property.set_maximum_retry_count=15

The default value for this property is five (5). To ensure that the Hub tries to establish the network connection enough times to account for network traffic, set this property to a high number (such as 15).

58 Chapter 5: Trillium Director Cleanse Engine

Page 59: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

C H A P T E R 6

SAP Data Services XI Cleanse Engine

This chapter includes the following topics:

• SAP Data Services XI Cleanse Engine Overview, 59

• About SAP Data Services XI Integration, 59

• Adding an SAP Library in the Cleanse Functions Tool, 61

• Configuring Generated Libraries, 64

SAP Data Services XI Cleanse Engine OverviewOne of the ways to access data cleansing functionality in the SAP Data Services XI product is through Web services that the SAP product publishes. This chapter describes how, in your Informatica MDM Hub implementation, to set up and use the Informatica MDM Cleanse Adapter for SAP (Web Services) to access cleanse functions that are published as Web services.

This functionality allows you add a new type of cleanse library - an SAP cleanse library - to your Informatica MDM Hub implementation, and then integrate cleanse functions in the SAP library into your mappings, just as you would integrate any other type of cleanse function available in your Informatica MDM Hub implementation. Informatica MDM Hub acts as a Web service client application that consumes Web services published by SAP Data Services XI.

If you do not use the SAP Data Services XI cleanse engine, you can skip this chapter.

About SAP Data Services XI IntegrationYou must configure the Informatica MDM Hub implementation, to set up and use the Informatica MDM Cleanse Adapter for SAP (Web Services) to access cleanse functions that are published as Web services.

PrerequisitesTo use this functionality, you must have installed the following software:

• Siperian Hub XU SP2 Patch B or later

59

Page 60: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

• SAP BusinessObjects Data Services XI 3.2 or higher

• Informatica MDM license that enables SAP Data Services XI functionality (siperian.sap_=yes in the configuration file)

For more information, see the Informatica MDM Hub Configuration Guide . For more information about SAP BusinessObjects Data Services XI 3.2, see the following URL: http://help.sap.com/businessobject/product_guides/boexir32/en/xi32_ds_doc_map_en.pdf

Process OverviewTo use these Web services in your Informatica MDM Hub implementation, complete the following steps:

1. Obtain a SAP Web Services Description Language (WSDL) file for the Web services that you want to consume. Use the WSDL generated by SAP BusinessObjects Data Services XI 3.2. You must add to the WSDL any Webservices that you want to consume.

2. In the Cleanse Functions tool of the Informatica MDM Hub Console, add an SAP library and specify the URI of the Web Services Description Language (WSDL) file as well as connection information (service and port) to hosted Web services. For instructions, see “Adding an SAP Library in the Cleanse Functions Tool” later in this chapter.

The Cleanse Functions tool builds the SAP library based on the WSDL, and then displays the list of cleanse functions defined in the WSDL. Each cleanse function represents a separate Web service.

3. Enable a generated library according to the instructions in “Enabling a Generated Library” later in this chapter.

4. In the Mappings tool of the Informatica MDM Hub Console, use the available cleanse functions in your mappings as required. You must configure the inputs and outputs just as you would configure any other mappings.

For more information, see the Informatica MDM Hub Configuration Guide .

60 Chapter 6: SAP Data Services XI Cleanse Engine

Page 61: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Run-time BehaviorAt run time, each time a cleanse function is fired:

• Informatica MDM Hub submits the request to the associated Web service, passing the input parameters from the currently-processed record as defined in the mapping.

• Informatica MDM Hub handles the response from the Web service, retrieving output parameters as defined in the mapping, and writing the appropriate changes to the record.

Considerations• Web service invocations are synchronous only. Asynchronous invocations are not supported.

• Web service invocations operate on a single record at a time. Batch processing of multiple records in a single service invocation is not supported.

• The purpose of using the Web services is strictly to transform data that is passed in the request according to the associated cleanse function. Other types of Web services, such as publish/subscribe services, are not supported.

• If the Web service returns an error, Informatica MDM Hub moves the record to the reject table and saves a description of the problem (including any error information returned from the Web service).

• If the Web service is published on a remote system, the infrastructure must be in place for Informatica MDM Hub to connect to the Web service (such as a network that accesses the Internet).

• When using cleanse functions that are implemented as Web services, run-time performance of Web service invocations depends on some factors that are external to Informatica MDM Hub, such as availability of the Web service, the time required for the Web service to process the request and return the response, and network speed.

• You can run WSDL cleanse function with multi-threading. To enable this, change the thread count on the Informatica MDM Hub Process Server. Ensure that there is a sufficient number of instances of the SAP Web Service to handle the multiple Informatica MDM cleanse threads; otherwise, records might be rejected due to timeouts.

• WSDL files must comply with the Axis2 Databinding Framework (ADB). Non-compliant WSDL files are not supported.

• When you configure mappings, you must ensure that the inputs and outputs are appropriate for the Web service you are calling. The Mappings tool does not validate your inputs and outputs – this is done by the Web service instead. If you have invalid inputs or outputs, the Web service returns an error response and processed records are moved to the reject table with an explanation of the error.

• The Cleanse Functions tool displays cleanse function parameters but nodocumentation, if specified, in the WSDL file (embedded <annotation> or <documentation> tags).

Adding an SAP Library in the Cleanse Functions ToolOnce you have installed the prerequisite software and obtained an SAP WSDL file, you use the Cleanse Functions tool in the Informatica MDM Hub Console to add the SAP library to your Informatica MDM Hub implementation.

1. Launch the Informatica MDM Hub Console, if it is not already running.

2. Start the Cleanse Functions tool.

3. Obtain a write lock (Write Lock > Acquire Lock).

Adding an SAP Library in the Cleanse Functions Tool 61

Page 62: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

4. Select the Cleanse Functions (root) node.

5. Choose Cleanse Functions > Add SAP Data Services XI Library.

6. In the Add SAP Library dialog, specify the following settings:

Setting Description

Library Name Name of this SAP Data Services XI library. You can assign any arbitrary name that helps you classify and organize the collection of SAP cleanse functions.Consider including SAP in the name to distinguish this from other cleanse function libraries. This name appears as the folder name in the Cleanse Functions list.

SAP Data Services XI WSDL URI

URI (location) of the SAP WSDL to implement.Note: URI paths which includes spaces (e.g C:\Documents and Settings\) are not permitted. Also, ensure the path separators entered are appropriate for the operating system.

SAP Data Services XI WSDL Service

Service of the SAP WSDL to implement.

SAP Data Services XI WSDL Port

Port of the SAP WSDL to implement.

Description Descriptive text for this library that you want displayed in the Cleanse Functions tool.

Note: Simple WSDLs often have only one Service and one Port. You can refer to the SAP WSDL for the values to specify for these settings.

The following figure shows sample settings for an SAP WSDL that invokes default SAP name and address cleansing.

7. Click OK to add the metadata definition for this new SAP library to the local ORS repository.

62 Chapter 6: SAP Data Services XI Cleanse Engine

Page 63: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

8. Click the Refresh button to generate the SAP library. The Cleanse Functions tool retrieves the latest SAP WSDL, generates the SAP library, and displays any available cleanse functions in the Cleanse Functions list.

Note:

• If the Cleanse Functions tool cannot consume the SAP WSDL file (for example, due to a syntax error), then it displays an error message instead. You must fix the SAP WSDL file or obtain a valid one.

• The maxOccurs attribute in the WSDL file must be set to '1'.

• If changes in the SAP WSDL file affect any existing mappings, the Cleanse Functions tool displays an error. You must fix the affected mappings before running the cleanse process.

9. Click a cleanse function to display its properties.

10. At this point, you can add these cleanse functions to your mappings in the Mappings tool, as shown in the “Process Overview” earlier in this chapter.

11. Test the function by clicking the Test tab and then clicking the Test button.

Adding an SAP Library in the Cleanse Functions Tool 63

Page 64: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Configuring Generated LibrariesWhen the Cleanse Functions tool generates a library, it creates a set of properties files in the SIP_HOME/cleanse/lib directory with the following naming format:

siperian-cleanse-<servicename>_<functionname>.properties.new

Note: File names must be unique (service name + function name) within this directory.

Enabling a Generated LibraryTo enable a generated library, remove the .new extension from the file name.

Properties File Syntax# endPointOverride =## Inputs for <functionname>##<wsdl_in_param1> = name , description#<wsdl_in_param2> = name , description## Outputs for <functionname>##<wsdl_out_param1> = name , description#<wsdl_out_param2> = name , description

Example Properties FileBelow is an example code listing.

#endPointOverride =## Inputs for name_Address#Req_Name_DITypeVarchar1024 = Req_Name_DITypeVarchar1024 , Req_Name_DITypeVarchar1024Req_State_DITypeVarchar1024 = Req_State_DITypeVarchar1024 ,Req_State_DITypeVarchar1024Req_Address_DITypeVarchar1024 = Req_Address_DITypeVarchar1024 ,Req_Address_DITypeVarchar1024Req_City_DITypeVarchar1024 = Req_City_DITypeVarchar1024 , Req_City_DITypeVarchar1024Req_Pcode_DITypeVarchar1024 = Req_Pcode_DITypeVarchar1024 ,Req_Pcode_DITypeVarchar1024#

64 Chapter 6: SAP Data Services XI Cleanse Engine

Page 65: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

# Outputs for name_Address#Res_Address_DITypeVarchar1024 = Res_Address_DITypeVarchar1024 ,Res_Address_DITypeVarchar1024

Res_Name_DITypeVarchar1024 = Res_Name_DITypeVarchar1024 , Res_Name_DITypeVarchar1024Res_State_DITypeVarchar1024 = Res_State_DITypeVarchar1024 ,Res_State_DITypeVarchar1024Res_City_DITypeVarchar1024 = Res_City_DITypeVarchar1024 , Res_City_DITypeVarchar1024Res_Pcode_DITypeVarchar1024 = Res_Pcode_DITypeVarchar1024 ,Res_Pcode_DITypeVarchar1024

Editing a Generated Properties FileTo edit a generated properties file:

1. Remove the .new from the end of the file name to enable it.

2. Open the file in a text editor.

3. Make the changes you want to the file, then save the file.

4. In the Cleanse Functions tool, select the SAP library, then click the Refresh button to activate these changes in the properties file.

Note: When you click Refresh to generate a library from a WSDL, the Cleanse Functions tool regenerates the file with the .new name. Changes made to a file that has been renamed (.new was removed) is not be overwritten when the library is refreshed.

You would edit a generated properties file for the following reasons:

• to change the connection endpoint

• to rename or remove extraneous parameters

Adding a Different Connection EndpointIf the communication endpoint for a Web service changes, or if you must point to a different environment, you can edit the properties file and provide the alternate endpoint URL as an override:

# endPointOverride =This setting overrides the configured endpoint setting for this SAP library.

You might need to do this, for example, when moving from a development/test environment to a production environment, or when (in a production environment) it might not be feasible to perform a refresh on the adapter.

Configuring Generated Libraries 65

Page 66: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Managing Input and Output ParametersThe names generated from the WSDL can sometimes be long and difficult to read, or in an uncommon order, as shown in the following example.

You can edit the generated parameter file and rename parameters to make them easier to read and recognize. In fact, if you encounter an error indicating that one or more names are too long to be stored in the ORS repository (> 100 characters), then you must shorten these names.

You can simplify the library by removing extraneous parameters that are not needed for your Web service invocations. You can also reorder parameters.

To edit parameter names:

• For required parameters (that you want to use in your cleanse functions), uncomment by removing the # character. Any parameter that is not uncommented is removed from the function.

• Edit parameter names and descriptions as required. You cannot have duplicate parameter names in one function (duplicates are ignored).

The library must be refreshed for the changes to the properties file to take effect. In the Cleanse Functions tool, select the library and click Refresh.

66 Chapter 6: SAP Data Services XI Cleanse Engine

Page 67: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

For example, the following figure shows the default generated output for an SAP library.

To rename a parameter, change its name setting:

<parameter> = name , descriptionThe following example shows renamed parameters.

In addition, you can change the order of the parameters, as shown in the following example.

Configuring Generated Libraries 67

Page 68: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Save changes to the file and then, in the Cleanse Functions tool, select the library and click Refresh to update the modified cleanse function settings. The changes are reflected in Function Inputs and Function Outputs.

Once refreshed, be sure to test the cleanse function to verify its operation.

68 Chapter 6: SAP Data Services XI Cleanse Engine

Page 69: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

C H A P T E R 7

TroubleshootingThis chapter includes the following topics:

• AddressDoctor Initialization Issues, 69

• Cleanse Engine Initialization Fails, 69

• Remote Initialization Fails, 70

• Trillium Errors, 70

• Initialization Fails for MDM Console, 70

• Trillium Director Integration, 71

AddressDoctor Initialization IssuesIf you cannot initialize AddressDoctor after you install the cleanse match server, use the following suggestions to troubleshoot the issue:

1. Check the cmxserver.log file to confirm that AddressDoctor initialization failed.

2. Ensure that your Informatica MDM license has AddressDoctor enabled. If it does not, contact Informatica Support. When you obtain the correct license, restart your application server and begin the installation again.

3. Try rebooting your machine (Windows) or refreshing your environment (UNIX).

4. UNIX: Ensure that the SSAPR environment variable has been added to your environment, and that it points to the current Informatica MDM files. Restart the application server.

5. Ensure you have the correct jade.dll file (Windows) or libJADE.so file (UNIX) and that it is the correct size, then restart the application server.

Note: Check the system path for older jade.dll files and delete them. Accumulation of these older files is a common occurrence in development environments.

Cleanse Engine Initialization FailsThis section discusses workarounds and tips for initialization failures.

69

Page 70: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

In general, cleanse issues fall into the following categories:

• Issues that occur during the application server initialization:

- Lack of a license file.

- Erroneous or missing cmxcleanse.properties file entries.

- Errors when loading system files.

• Issues that occur when the Hub is making calls to the cleanse tool in batch or real time:

- Uninitialized adapter.

- Port conflicts.

- No response from the cleanse tool server.

Remote Initialization FailsIf your remote initialization (initializing cleanse engines that are on a remote machine) fails, use the following suggestions to troubleshoot the issue:

• Ensure that the cleanse engines are running, especially if the cleanse engine you are using has a service associated with it.

• Check that your properties file already points to the specific machine that is running that cleanse engines.

• Restart the application server.

• Ensure the remote machine running the cleanse engine services is started before the application server is turned on.

Trillium ErrorsIn general, Trillium errors fall into the following categories:

• Lack of read, write, execute permissions

• Port conflicts

• Missing Trillium Director client library TGenClient.lib. This library should exist on the same machine that is running the Process Server.

Initialization Fails for MDM ConsoleIf your Trillium initialization fails for the MDM Console, (such as when you get the following error: Cannot borrow data from the POOL), try stopping the application server and the Trillium Director services. Restart the Trillium Director services first, then the application server.

Note: See the Informatica Knowledge Base (https://communities.informatica.com) for additional information.

70 Chapter 7: Troubleshooting

Page 71: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Remote Initialization FailsIf your remote initialization (initializing cleanse engines that are on a remote machine) fails, use the following suggestions to troubleshoot the issue:

• Ensure that the cleanse engines are running, especially if the cleanse engine you are using has a service associated with it.

• Check that your properties file already points to the specific machine that is running that cleanse engines.

• Restart the application server.

• Ensure the remote machine running the cleanse engine services is started before the application server is turned on.

Trillium Director IntegrationThe primary tool for troubleshooting your Trillium Director integration is the working files written by Trillium Director. You can also check the Process Server files.

About Trillium Director Working FilesWhen a batch stage job is executed, the Process Server saves working files. When a Trillium Director function is used, a Trillium Director file is created for the purpose of auditing and troubleshooting.

Finding the Location of the Trillium Director Work FilesTrillium Director writes logs and work files to the logs directory and other places defined by the project. By default, the logs directory is:

• Windows:C:\tril7v8\Director\logs

• UNIX:/u1/tril7v8/Director/logs

To find the location of the work files, look at the following parameter in the properties file:

cmx.server.datalayer.cleanse.working_files.location

Setting the Location of the Process Server Work FilesProcess Server writes working files to a directory specified in the cmxcleanse.properties file.

To set the location of the work files, set the following parameter in the properties file:

cmx.server.datalayer.cleanse.working_files.location

Setting Whether Working Files are KeptTo set whether work files are kept, set the following parameter in the cmxcleanse.properties file:

cmx.server.datalayer.cleanse.working_files=KEEPSet this property value to FALSE to discard these working files. If you specify that the working files are not kept, the value for the location of the files is irrelevant.

Trillium Director Integration 71

Page 72: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Setting the Number of Connections in a Connection PoolWhen you use large data sets with the Trillium adapter, the adapter uses a pool of connections, which is shared by all jobs in an application server.

The following parameter controls the pool level:

cleanse.library.trilliumDir.property.td.project.pool.MaxActive=nnBy default, the adapter allows as many connection attempts as are requested. However, you can set this parameter if you must restrict the connections.

72 Chapter 7: Troubleshooting

Page 73: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Index

AAddressDoctor cleanse engine

configuring 23configuring JVM settings 45initialization issues 69required files 24

upgrade editing the properties file 26required files 26

upgrading 26

BBusiness Objects DQ XI cleanse engine

configuring 48Informatica MDM Hub integrating with 48

Ccleanse engine installation

prerequisites 10cleanse engines

initialization failure 69supported 10

Cleanse Match Server AddressDoctor cleanse engine 23FirstLogic cleanse engine 47IDQ cleanse engine 12

cmxcleanse.properties file editing 24

connection pool, setting connections 72

FFirstLogic cleanse engine

components installation 48configuring to use with Informatica MDM Hub 48integration 47library 49transactional mode 49transactional mode sample 49

HHP-UX

JVM settings 46Hub Server

parameters, updating 70, 71

IIDQ cleanse engine

Adding IDQ Library 15integration 12

Informatica MDM Hub configuring to use with FirstLogic 48

installation troubleshooting 69

JJBoss

JVM settings on HP-UX and Solaris 46JVM settings on Windows 46

JVM settings configuring for AddressDoctor 45JBoss on HP-UX and Solaris 46JBoss on Windows 46WebLogic on HP-UX and Solaris 46WebLogic on Windows 45WebSphere on Windows 45

Llicense requirements 47

MMDM Console

initialization failures 70

Nnetwork connection retries, setting 58

Ppreface 7prerequisites

cleanse engine installation 10properties file

editing 24editing a generated file 18

SSAP Data Services XI cleanse engine

integration 59overview 59

73

Page 74: Cleanse Adapter Guide - Informatica Documentation/4/MDM… · 1 • • • • • • ...

Solaris JVM settings 46

Tthread-pool connections, setting 58Trillium

errors 70troubleshooting integration 71

Trillium Director cleanse engine displaying in console 52integration 51multithreading 57

Trillium Director cleanse engine (continued)prerequisites 52sample configuration file 54td11_default_config_Global.txt file 54td11_default_config_US_detail.txt file 54td11_default_config_US_summary.txt file 54testing configuration 54troubleshooting 71upgrading 56work files 71

Trillium patches 51troubleshooting

installation 69

74 Index


Recommended