• How to extract text from PDF(Image) files, OCR


    Background: below is SS1.0 as example since it came from NetSuite email plugin, SS2.0 is the same thing.

    1. Registry a API key throw https://ocr.space/OCRAPI

    There are limitations for Free Plan

    2. Save the email attachment(PDF file) to NetSuite FileCabinet, set it to available without login, get the full url address, encode it.

    var importFile = attachments[indexAtt];importFile.setIsOnline(true);
    var intFileId = nlapiSubmitFile(importFile);
    var strInvFileUrl = "https://" + nlapiGetContext().getCompany() + ".app.netsuite.com"+ objInvoiceFileRec.getURL();
    strInvFileUrl = encodeURIComponent(strInvFileUrl);

     

    3. Send Request to https://api.ocr.space/parse/imageurl?apikey=abcAPIKEYabc&filetype=PDF&isTable=true&url=

    var response = nlapiRequestURL(strReqUrl, null, a);
    There are varience of parameters for this API, in my case, it's invoice formated as table, that's why I send isTable=true to identify it; then it will help me to locate the expected cell and values.


    4. Got and parsed the Response, we will get the Text messages on the PDF or Images.

    var arrParsedLines = (objOcrRes['ParsedResults'] && objOcrRes['ParsedResults'][0]) ? objOcrRes['ParsedResults'][0]['TextOverlay']['Lines']: null;
    var objVndBillData = parseDataFromInvPdf(arrParsedLines);

  • 相关阅读:
    Weex-语法笔记 一
    Entitlements (授权机制) 延伸
    FOUNDATION_EXPORT 延伸
    NS_ASSUME_NONNULL_BEGIN 延伸
    SSLPinning 延伸
    Socket.io 延伸
    UI_APPEARANCE_SELECTOR 延伸
    Modules和Autolinking
    代码块(Block)回调一般阐述
    数据库连接池配置错误导致OOM
  • 原文地址:https://www.cnblogs.com/backuper/p/How_to_extract_text_from_PDF_or_Image_files.html
Copyright © 2020-2023  润新知