public class WikipediaTokenizer extends Tokenizer
Modifier and Type | Field and Description |
---|---|
static int |
ACRONYM_ID |
static int |
ALPHANUM_ID |
static int |
APOSTROPHE_ID |
static java.lang.String |
BOLD |
static int |
BOLD_ID |
static java.lang.String |
BOLD_ITALICS |
static int |
BOLD_ITALICS_ID |
static int |
BOTH |
static java.lang.String |
CATEGORY |
static int |
CATEGORY_ID |
static java.lang.String |
CITATION |
static int |
CITATION_ID |
static int |
CJ_ID |
static int |
COMPANY_ID |
static int |
EMAIL_ID |
static java.lang.String |
EXTERNAL_LINK |
static int |
EXTERNAL_LINK_ID |
static java.lang.String |
EXTERNAL_LINK_URL |
static int |
EXTERNAL_LINK_URL_ID |
static java.lang.String |
HEADING |
static int |
HEADING_ID |
static int |
HOST_ID |
static java.lang.String |
INTERNAL_LINK |
static int |
INTERNAL_LINK_ID |
static java.lang.String |
ITALICS |
static int |
ITALICS_ID |
static int |
NUM_ID |
static java.lang.String |
SUB_HEADING |
static int |
SUB_HEADING_ID |
static java.lang.String[] |
TOKEN_TYPES
String token types that correspond to token type int constants
|
static java.lang.String[] |
tokenImage
Deprecated.
Please use
TOKEN_TYPES instead |
static int |
TOKENS_ONLY |
static int |
UNTOKENIZED_ONLY |
static int |
UNTOKENIZED_TOKEN_FLAG
This flag is used to indicate that the produced "Token" would, if
TOKENS_ONLY was used, produce multiple tokens. |
Constructor and Description |
---|
WikipediaTokenizer(java.io.Reader input)
Creates a new instance of the
WikipediaTokenizer . |
WikipediaTokenizer(java.io.Reader input,
int tokenOutput,
java.util.Set untokenizedTypes) |
public static final java.lang.String INTERNAL_LINK
public static final java.lang.String EXTERNAL_LINK
public static final java.lang.String EXTERNAL_LINK_URL
public static final java.lang.String CITATION
public static final java.lang.String CATEGORY
public static final java.lang.String BOLD
public static final java.lang.String ITALICS
public static final java.lang.String BOLD_ITALICS
public static final java.lang.String HEADING
public static final java.lang.String SUB_HEADING
public static final int ALPHANUM_ID
public static final int APOSTROPHE_ID
public static final int ACRONYM_ID
public static final int COMPANY_ID
public static final int EMAIL_ID
public static final int HOST_ID
public static final int NUM_ID
public static final int CJ_ID
public static final int INTERNAL_LINK_ID
public static final int EXTERNAL_LINK_ID
public static final int CITATION_ID
public static final int CATEGORY_ID
public static final int BOLD_ID
public static final int ITALICS_ID
public static final int BOLD_ITALICS_ID
public static final int HEADING_ID
public static final int SUB_HEADING_ID
public static final int EXTERNAL_LINK_URL_ID
public static final java.lang.String[] TOKEN_TYPES
public static final java.lang.String[] tokenImage
TOKEN_TYPES
insteadpublic static final int TOKENS_ONLY
public static final int UNTOKENIZED_ONLY
public static final int BOTH
public static final int UNTOKENIZED_TOKEN_FLAG
TOKENS_ONLY
was used, produce multiple tokens.public WikipediaTokenizer(java.io.Reader input)
WikipediaTokenizer
. Attaches the
input
to a newly created JFlex scanner.input
- The Input Readerpublic WikipediaTokenizer(java.io.Reader input, int tokenOutput, java.util.Set untokenizedTypes)
public Token next(Token reusableToken) throws java.io.IOException
next
in class TokenStream
java.io.IOException
public void reset() throws java.io.IOException
reset
in class TokenStream
java.io.IOException
Copyright © 2000-2014 Apache Software Foundation. All Rights Reserved.