Using Python to Call the Baidu Translation API to Batch-Translate Office, JSON, and MySQL Data

Request an API Key

Step 1: Open Baidu Translate Open Platform Go to the official website and sign up for a Baidu account.
Step 2: Go to the Management Console, select Developer Center > Developer Information, and complete the personal verification under "Basic Information."

Step 3: Create an API. Go to the Developer Center > API Key Management and create an API.

Call this API using Python

The key is to write a `translate` function: generate a random number, concatenate the string, calculate the MD5 hash, send a GET request, and check `trans_result`. The code is below.

import requests

import hashlib

import random

import os

import time

# is read from an environment variable to avoid hard-coding the key in the code.

APP_ID = os.environ.get("BAIDU_APP_ID", "Your APP_ID")

SECRET_KEY = os.environ.get("BAIDU_SECRET_KEY", "Your secret key")

def translate(text, from_lang="auto", to_lang="en"):

"""

Call the Baidu Translate API to translate a piece of text.

text: String to be translated

from_lang: Source language code; "auto" indicates automatic detection

to_lang: Target language code, such as en, jp, fra

Returns: The translated string; returns `None` on failure

"""

url = "https://fanyi-api.baidu.com/api/trans/vip/translate"

# Authenticated Premium Edition has a limit of 10 requests per second, so wait 0.1 seconds before each request.

time.sleep(0.1)

# generates random numbers to prevent requests from being cached

salt = str(random.randint(32768, 65536))

# Concatenate the strings in the order specified by Baidu, then calculate the MD5 hash.

sign_str = APP_ID + text + salt + SECRET_KEY

sign = hashlib.md5(sign_str.encode("utf-8")).hexdigest()

# Assembly Request Parameters

params = {

"q": text,

"from": from_lang,

"to": to_lang,

"appid": APP_ID,

"salt": salt,

"sign": sign

}

# sends a GET request

response = requests.get(url, params=params, timeout=10)

result = response.json()

# Check the return result

if "trans_result" is in result:

#: Baidu returns a list, with each element being a translation result.

return result["trans_result"][0]["dst"]

else:

# displays error messages to facilitate troubleshooting

print(f"Translation failed: {result}")

return None

How to Choose a Target Language

Baidu Translate uses short codes to represent languages. Common ones include:zh Chinese,en English,jp Japanese,kor Korean,fra French,spa Spanish,de German,ru Russian,auto Indicates automatic detection of the source language. When calling the function, simply swap `from_lang` and `to_lang`; for example, to translate from Chinese to English: trans(text, from_lang="zh", to_lang="en"), the translation from English to Japanese is translate(text, from_lang="en", to_lang="jp"). Baidu Translate supports a wide range of languages; for a complete list, please refer to the official documentation. Here, we’ve listed only the most commonly used ones.

Translate Office Files

There are three types of Office files: Word, Excel, and PowerPoint. Python has corresponding libraries for reading and writing each of them. These three file types have different structures and are handled differently.
I use Word documents python-docx. The text is scattered throughout paragraphs and tables, so you have to go through every paragraph one by one, send each section for translation, and then paste the translation back into its original location.

from docx import Document

def translate_word(input_path, output_path, to_lang="en"):

"""

Translate all paragraphs and table text in the Word document.

input_path: Source file path

output_path: Path where the translated files are saved

to_lang: Target Language

"""

doc = Document(input_path)

# Iterate through all paragraphs, skipping empty paragraphs

for para in doc.paragraphs:

if para.text.strip():

translated = translate(para.text, to_lang=to_lang)

if translated:

# Clear the original paragraph and insert the translation

para.clear()

para.add_run(translated)

# Iterate through all tables and translate each cell

for table in doc.tables:

for row in table.rows:

for cell in row.cells:

if cell.text.strip():

translated = translate(cell.text, to_lang=to_lang)

if translated:

cell.text = translated

doc.save(output_path)

print(f"Translation complete: {output_path}")

I use Excel openpyxl. The key is the cells; you need to go through each cell that contains a value one by one, translate it, and then write it back.

import openpyxl

def translate_excel(input_path, output_path, to_lang="en"):

"""

Translate the text in all cells of the Excel file.

Translate only cells containing text; leave numbers and formulas unchanged.

"""

wb = openpyxl.load_workbook(input_path)

# Iterate through every worksheet

for sheet in wb.worksheets:

#: Iterate through all rows and cells

for row in sheet.iter_rows():

for cell in row:

# processes only string content and skips numbers, dates, etc.

if cell.value and isinstance(cell.value, str):

translated = translate(cell.value, to_lang=to_lang)

if translated:

cell.value = translated

wb.save(output_path)

print(f"Translation complete: {output_path}")

I use PowerPoint python-pptx. The text is scattered throughout text boxes and shapes on each page, so you'll need to go through them layer by layer.

from pptx import Presentation

def translate_ppt(input_path, output_path, to_lang="en"):

"""

Translate the text in all text boxes in the PowerPoint presentation.

Iterate through each slide, then iterate through each text box in each shape.

"""

prs = Presentation(input_path)

for each slide in prs.slides:

for shape in slide.shapes:

# Check if this shape contains a text box

if shape.has_text_frame:

for para in shape.text_frame.paragraphs:

if para.text.strip():

translated = translate(para.text, to_lang=to_lang)

if translated:

# Replace the entire paragraph directly

para.text = translated

prs.save(output_path)

print(f"Translation complete: {output_path}")

By combining these three functions and adding logic to identify file types based on their extensions, the program can batch-process all Office files in a folder.

Translate JSON Files

JSON has a nested structure: a key can have a value that is either a string or an object, and that object may contain additional key-value pairs. I use a recursive function to extract all string values, translate them, and then put them back.
The logic behind the recursion is as follows: if it encounters a dictionary, it iterates through its values; if it encounters a list, it iterates through its elements; if it encounters a string, it translates it; and if it encounters any other type, it returns it as-is. This recursive function can handle JSON no matter how deeply nested it is.

import json

def translate_json_value(obj, to_lang="en"):

"""

Recursively traverse the JSON data and translate all string values.

Dictionaries, lists, and strings are all handled separately.

Non-string types (numbers, booleans, null) are returned as-is.

"""

if isinstance(obj, dict):

# If it's a dictionary, iterate through each key-value pair

return {key: translate_json_value(value, to_lang) for key, value in obj.items()}

elif isinstance(obj, list):

# If it's a list, iterate through each element

return [translate_json_value(item, to_lang) for item in obj]

elif isinstance(obj, str):

# If it's a string, translate it

if obj.strip():

translated = translate(obj, to_lang=to_lang)

return translated if translated, else obj

return obj

else:

# Directly returns numeric, Boolean, and null values

return obj

def translate_json_file(input_path, output_path, to_lang="en"):

"""

Read a JSON file, translate all string values, and write the results to a new file.

"""

with open(input_path, "r", encoding="utf-8") as f:

data = json.load(f)

translated_data = translate_json_value(data, to_lang)

with open(output_path, "w", encoding="utf-8") as f:

# ensure_ascii=False ensures that Chinese characters are displayed correctly and are not converted to \uXXXX

json.dump(translated_data, f, ensure_ascii=False, indent=2)

print(f"Translation complete: {output_path}")

Translating Fields in a MySQL Database

Databases work differently from files. Files are treated as a single unit, while databases are organized row by row. You need to connect to the database, retrieve the content to be translated, translate it, and then write it back. First, install the dependencies:pip install pymysql requests。
The logic here is straightforward: first look it up, then translate it, and then update it. When linking, use charset="utf8mb4"Otherwise, the database may not be able to store Chinese characters and emojis. Enclose table and column names in backticks to prevent conflicts with SQL keywords.

import pymysql

def translate_mysql_table(host, user, password, db_name, table_name, column_name, to_lang="en"):

"""

Translate all text values in a specific column of a MySQL table.

host/user/password/db_name: Database connection information

table_name: The name of the table to be translated

column_name: The column name to be translated

to_lang: Target Language

"""

# Connecting to the Database

conn = pymysql.connect(

host=host, user=user, password=password,

db=db_name, charset="utf8mb4"

)

cursor = conn.cursor()

# Retrieve all non-null values, including the primary key `id`, to facilitate future updates.

sql_select = f"SELECT id, `{column_name}` FROM `{table_name}` WHERE `{column_name}` IS NOT NULL AND `{column_name}` != ''"

cursor.execute(sql_select)

rows = cursor.fetchall()

print(f"A total of {len(rows)} records to be translated were found")

# Translate and update line by line

for row_id, original_text in rows:

translated = translate(original_text, to_lang=to_lang)

if translated:

sql_update = f"UPDATE `{table_name}` SET `{column_name}` = %s WHERE id = %s"

cursor.execute(sql_update, (translated, row_id))

# Commit all changes

conn.commit()

cursor.close()

conn.close()

print("Translation complete")

Previous Article Recommended Image Compression Websites: These 6 Are Definitely Worth Bookmarking